From ef93154b6528054466744df0eec2f26bf1575002 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 17:49:08 -0400 Subject: [PATCH 001/108] feat(mq): give every tenant a queue of its own --- AGENTS.md | 8 +- CHANGELOG.md | 8 +- clients/ts/src/dlq.ts | 18 +- clients/ts/src/namespaces.test.ts | 19 + clients/ts/src/types.ts | 4 +- docs/src/content/docs/api.md | 9 +- docs/src/content/docs/architecture.md | 16 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 18 +- docs/src/content/docs/durability.md | 6 +- docs/src/content/docs/ingest-pipeline.md | 26 +- docs/src/content/docs/sdk/admin.md | 7 + docs/src/content/docs/sdk/reference.md | 4 +- docs/src/content/docs/settings-directory.mdx | 16 +- docs/src/content/docs/why-wavehouse.md | 8 +- internal/api/dlq.go | 29 +- internal/api/dlq_test.go | 138 ++-- internal/api/ingest.go | 2 +- internal/api/router_test.go | 5 +- internal/app/app.go | 9 +- internal/app/app_test.go | 56 +- internal/app/wire.go | 117 +-- internal/ingest/sweeper.go | 41 +- internal/ingest/sweeper_test.go | 31 +- internal/ingest/worker.go | 19 +- internal/ingest/worker_test.go | 28 +- internal/mq/embedded.go | 756 ++++++++++++++----- internal/mq/embedded_test.go | 649 ++++++++++++---- internal/mq/mq.go | 130 ++-- internal/mq/subject.go | 76 +- internal/mq/subject_test.go | 52 +- internal/settings/settings.go | 23 +- internal/settings/store.go | 2 +- internal/stream/subscriber.go | 20 +- internal/testutil/mocks.go | 16 +- internal/testutil/testutil.go | 19 + 36 files changed, 1679 insertions(+), 708 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 59f08ef3..16595721 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive`/`longestGapWindow` for the two settings folded over every tenant served, and `defaultSetting`/`onDefaultAdopt` for the one resource a process still has one of, the MQ, which follows tenant `0`; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `...
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config @@ -38,14 +38,14 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, a topic without one is refused, and a pre-tenant subject reads as tenant `0`'s) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds the byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload) and `Stats` (the system gauges' source). Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` +- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) - **`policy/`** — Hasura-style access control, **role-first**: `TablePolicy` is `map[string]RolePermissions`, and a role's grant splits by operation into `SelectPermissions` (columns, row `filter`, aggregations, the `max_*` limits) and `InsertPermissions` (columns, `check`) — so a field only one side honors does not exist on the other. `Evaluate()` resolves ONE operation and leaves the other side **nil** (`Select *ResolvedSelect` / `Insert *ResolvedInsert`), which every accessor fails closed on — nil is "not resolved", distinct from an empty side, which is "unrestricted" (what the admin return builds). Claim templating (`{{ jwt.claim.path }}`) resolves during that call. Policies come from `Source`, a `func() *Policy` read per call (`settings.Store.Policy` in production, `Static(p)` in tests) - **`query/`** — Structured query AST types + SQL builder with schema validation, structural policy predicate/limit emission, timestamp bucketing - **`settings/`** — the settings directory, in either shape ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): flat (the four files: tenant `0` alone) or nested (one folder per tenant, never mixed). `Validate` detects the shape and checks it — `ValidateDir` per directory (strict JSON, per-file rules, cross-file role references), folder names against `tenant.Parse`, a nested finding's `File` led by its folder; `Store` is a passive holder (one tenant's adopted snapshot, typed accessors read per call); `Registry` (tenant id → `Store`) owns `Open`, the serialized `Reload`/`ReloadTenant`, the `AfterAdopt` hooks, and the fsnotify `Watch` (flat only). Flat refuses an invalid directory at boot and keeps the previous snapshot on a rejected reload; nested fails closed per tenant (a rejected folder stops being served, the rest carry on, a whole-tree reload mirrors the folders, down to none, and a finding about the root itself rejects the reload whole). Plus the embedded (`go:embed`) seed `wavehouse bootstrap` writes - **`stream/`** — SSE fan-out: rows travel POSITIONALLY, so each connection is told its projected column list in an `event: schema` frame before its first row and again on drift — **not** guaranteed after a gap-fill across a column change, which can leave a connection reading live rows against a stale list until it reconnects ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)) — (tracked per connection; replay tracks its own). The event `Hub` (registers subscribers by `(mq.Topic, role)` — one tenant's table — and evaluates each event under its own tenant's policy and schema registry; `Prune` evicts the subscribers of every tenant a reload stopped serving; `Broadcast` projects + serializes each event once per role, the #294 delivery hot path — a role carrying a row-level `filter` keeps the shared projection but delivers per subscriber, each subscriber's claims evaluated against the row, #319), `Subscriber` (per-connection outbound `Frame` queue, `Send`/`Frames`; claims fixed at construction, immutable; `Evict` asks its handler to end the stream), the `Bucket` fan-out set (`subscriberSet`, one per `(topic, role)`), the `Heartbeater` keepalive wheel, and `Metrics` (the `wavehouse_sse_*` stream instruments) -- **`tenant/`** — the tenant identifier ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): `ID` (a validated string), `Parse` (letters, digits, `_`, `-`; ≤ 64 bytes — safe as a folder name and as an MQ subject token), `Default` (`"0"`), and `Header` (`X-Tenant-ID`). Imports nothing from the rest of the repo. `api.TenantMW` resolves the header against `settings.Registry` before auth on every `/v1` route outside `/v1/ops/*` (`400` malformed, `404` unknown, a bare `503` for a nested tenant whose folder was rejected) and puts the resolved `*settings.Store` in the request context; the ops routes that address one tenant (`GET /v1/ops/pipes[/{name}]`, `POST /v1/ops/settings/reload`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh`, `POST /v1/ops/query`) take a strictly parsed `?tenant=` instead; handlers read it once (`api.StoreFromContext`) and pass it down as an argument, and nothing below a handler reads context. The stream hub and the ingest worker read each message's tenant off its `mq.Topic` and their getters take it; the sweeper folds over the tenants served (`longestGapWindow`); each served tenant has a schema registry of its own (story 6) +- **`tenant/`** — the tenant identifier ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): `ID` (a validated string), `Parse` (letters, digits, `_`, `-`; ≤ 64 bytes — safe as a folder name and as an MQ subject token), `Default` (`"0"`), and `Header` (`X-Tenant-ID`). Imports nothing from the rest of the repo. `api.TenantMW` resolves the header against `settings.Registry` before auth on every `/v1` route outside `/v1/ops/*` (`400` malformed, `404` unknown, a bare `503` for a nested tenant whose folder was rejected) and puts the resolved `*settings.Store` in the request context; the ops routes that address one tenant (`GET /v1/ops/pipes[/{name}]`, `POST /v1/ops/settings/reload`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh`, `POST /v1/ops/query`, `GET /v1/ops/dlq/stats`) take a strictly parsed `?tenant=` instead; handlers read it once (`api.StoreFromContext`) and pass it down as an argument, and nothing below a handler reads context. The stream hub and the ingest worker read each message's tenant off its `mq.Topic` and their getters take it; the sweeper hands the MQ each served tenant's own gap window (`gapWindows`); each served tenant has a schema registry of its own (story 6) ## Key Design Decisions @@ -56,7 +56,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 3. **Schema-driven ingest** — `POST /v1/ingest?table={table}` takes flat JSON, validated against the discovered schema (unknown fields rejected, types/nullability enforced). No envelope. The **declared `Content-Type` chooses the format and the bytes never do** (arity within the JSON family is still the body's): no declaration, one whose **media type** is unsupported or unparseable, a comma-bearing value that, as a whole, does not parse as one media type, or repeated lines that **disagree**, is a `415` decided *before* the body is read. A malformed *parameter* on a comma-free line never costs the request (`; charset=a; charset=b` still reads as its media type), and repeated lines are accepted only when they all resolve to the same **supported** format — two agreeing `text/csv` lines are still a `415`. A body declared NDJSON stays NDJSON whatever its bytes, so a bad line is a per-record error rather than a silent re-framing; the reverse (NDJSON sent as `application/json`) is deliberately **not** caught — record one, `200`, the rest ignored ([#561](https://github.com/Wave-RF/WaveHouse/issues/561)). Fail-closed — preserve it when touching `internal/api`. 4. **Async ingestion** — ingest returns 200 after optional dedup + MQ publish; ClickHouse writes happen later via `StartIngestWorker`. NATS full → 503 + Retry-After. 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. -6. **Dead Letter Queue** — failed batch inserts publish to `WAVEHOUSE_DLQ` (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format` — a pre-v2 envelope carries none — or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. +6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format` — a pre-v2 envelope carries none — or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. 8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. diff --git a/CHANGELOG.md b/CHANGELOG.md index 5bf58019..23c0c715 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked to hold the ack floor that the one shared stream's purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. +- **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. @@ -24,7 +24,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Schema discovery captures each table's DDL, its columns' ordinals and default expressions, and the server version** (`internal/discovery/discovery.go`, `internal/testutil/testutil.go`): `Column` gains `DefaultExpression` and `Position` (both from a widened `system.columns` select), `TableSchema` gains `DDL` from `system.tables.create_table_query`, and `SchemaRegistry` gains `ServerVersion()` from a `SELECT version()` probe next to the existing `SELECT timezone()`. Groundwork for the native type layer, captured on the same refresh as the columns so a stale version cannot outlive the schemas it describes. That is a publication guarantee, not a same-server one: `chconn.Manager` resolves the connection per call, so a reload changing `clickhouse.addr` mid-refresh can still pair a version from one server with schemas from another — narrow, and self-correcting on the next refresh. `DDL` is `json:"-"` and does **not** appear in `/v1/ops/schema`: that endpoint marshals `TableSchema` straight to the client, and an external-engine table (S3, MySQL, PostgreSQL, Kafka) renders its wiring there unconditionally — endpoint, bucket or host, database, username, S3 access key id. ClickHouse masks the password itself as `[HIDDEN]` from ~23.9 (verified on 26.7.3), so the exposure is the topology rather than the secret — except on an older server, or one with `display_secrets_in_show_and_select` enabled. `position` and `default_expression` are additive fields in the response. A table listed in `system.tables` with no `system.columns` rows is skipped rather than published column-less, and both new queries fail the refresh on error exactly as `timezone()` and `system.columns` do — callers keep the prior cache and retry. -- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the `WAVEHOUSE` and `WAVEHOUSE_DLQ` stream limits in place via `EmbeddedNATS.Resize` — shrinking below the buffered size backpressures until the worker drains, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on `WAVEHOUSE_DLQ` and ack; off → leave it unacked for redelivery, never dropped; the DLQ stream and `GET /v1/ops/dlq/stats` now always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. +- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the worker drains, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. - **"Was this page helpful?" feedback widget on every docs page** (`docs/src/components/PageFeedback.astro` (new), `docs/src/components/Footer.astro`): a thumbs-up / thumbs-down vote below the page content, captured to PostHog as `docs_feedback` with `{ helpful, page }`. It renders from `Footer.astro`'s sidebar branch — the same indirection the Cloud CTA uses — rather than a per-page import or frontmatter flag, so every content page gets it automatically, including ones not written yet; it sits *below* the Cloud CTA on the pages that carry one, and splash pages (the homepage and 404) take the other footer branch and never render it. One vote per page per visitor: the choice is remembered in `localStorage` keyed by pathname, and a revisit renders the thanks message instead of re-prompting (storage is a nicety, not the record — a browser with storage disabled still votes). - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. @@ -32,7 +32,9 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). The sweeper keeps the longest `stream.gap_window_minutes` among the tenants being served: the ingest queue is one stream and a purge is one bound over it, so purging less is the safe direction until the streams are per tenant. `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. The subject change needs no drain of its own (the v2 envelope's drain, below, still applies): the durable consumers filter `ingest.>`, which the previous two-token subjects match, and a subject with no tenant token reads as tenant `0`'s, so the subject an event in flight arrived on changes nothing about how it is inserted, streamed, or parked, and rows already parked keep counting in `GET /v1/ops/dlq/stats` — which sums a table across tenants until the queue is per tenant; a gap-fill spanning the upgrade omits the pre-upgrade events for one gap window. Per-tenant JetStream streams, the DLQ shrink guard, and `?tenant=` on the DLQ stats route are story 5b, after a research spike. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/subscriber.go`, `internal/settings/{settings,store}.go`, `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch is shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes`, keeping no acknowledged history for a tenant no longer served, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. + +- **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. - **The query cache and its singleflight are keyed by tenant** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/settings/{store,registry}.go` (+ tests), `internal/api/{cache_key,pipes,structured_query}.go` (+ tests; `cache_tenant_test.go` new), `internal/ingest/worker.go` (+ tests), `docs/src/content/docs/{architecture,deployment,api,ingest-pipeline}.md`, `AGENTS.md`): story 8 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files — every key simply gains tenant `0`'s prefix. The tenant leads every key the cached read paths build: the key `POST /v1/query?table={table}` and `GET/POST /v1/pipes/{name}` cache a result under is `:query:` and is their singleflight key too, and a version namespace is `.
.
.` (`cache.Namespace` gains `Tenant`; scope stays where it was, and inert). So identical requests from two tenants are two entries and two flights to ClickHouse, a tenant is never served another tenant's cached rows, and a batch the ingest worker inserts bumps the namespaces of the tenant it was inserted for and no other's (`IngestWorker.invalidate` takes the tenant as a parameter — the worker's own until story 5 reads it off the message, so over a nested directory it is still tenant `0`'s namespaces every batch bumps, and another tenant's cached query results expire on their TTL alone until then). The pool stays one Ristretto instance sized by `cache.l1_max_cost`. The handlers read the tenant off the request's store — `settings.Store.Tenant`, stamped by the registry when it creates the store (the commit is shared with story 7) — so nothing new rides the request context. This lands ahead of story 6 on purpose: once each tenant has its own ClickHouse connection, a tenant-blind key would be a silent cross-tenant read. diff --git a/clients/ts/src/dlq.ts b/clients/ts/src/dlq.ts index a783e8eb..c1276225 100644 --- a/clients/ts/src/dlq.ts +++ b/clients/ts/src/dlq.ts @@ -1,7 +1,7 @@ import { err, ok } from "./errors.js"; -import { request } from "./http.js"; +import { request, tenantParam } from "./http.js"; import type { StreamController } from "./stream/controller.js"; -import type { DLQStats, HttpContext, Result, StreamOptions } from "./types.js"; +import type { DLQStats, HttpContext, OpsRequestOptions, Result, StreamOptions } from "./types.js"; type CreateStreamFn = (table: string, opts?: StreamOptions) => StreamController; @@ -15,23 +15,27 @@ export class DLQNamespace { this._createStream = createStream; } - /** Get DLQ statistics (message counts per table). */ - async list(opts?: { signal?: AbortSignal }): Promise> { + /** + * Get DLQ statistics (message counts per table) — of `opts.tenant`, the + * default tenant without it. A tenant with no dead-letter queue is a `404`. + */ + async list(opts?: OpsRequestOptions): Promise> { const { data, error } = await request(this._ctx, { method: "GET", path: "/v1/ops/dlq/stats", + params: tenantParam(opts), signal: opts?.signal, }); if (error) return err(error); return ok(data!); } - /** Get DLQ stats filtered by table name. */ - async table(name: string, opts?: { signal?: AbortSignal }): Promise> { + /** Get DLQ stats filtered by table name — of `opts.tenant`, the default tenant without it. */ + async table(name: string, opts?: OpsRequestOptions): Promise> { const { data, error } = await request(this._ctx, { method: "GET", path: "/v1/ops/dlq/stats", - params: { table: name }, + params: { table: name, ...tenantParam(opts) }, signal: opts?.signal, }); if (error) return err(error); diff --git a/clients/ts/src/namespaces.test.ts b/clients/ts/src/namespaces.test.ts index d1c0db5d..a32a9dfe 100644 --- a/clients/ts/src/namespaces.test.ts +++ b/clients/ts/src/namespaces.test.ts @@ -146,6 +146,25 @@ describe("DLQNamespace", () => { expect(fetchSpy.mock.calls[0][0]).toContain("table=clicks"); }); + it("list() and table() send opts.tenant as ?tenant=, and nothing without it", async () => { + fetchSpy.mockImplementation( + async () => new Response(JSON.stringify({ tables: {}, total: 0 }), { status: 200 }), + ); + const ns = new DLQNamespace(makeCtx(), mockStream); + + await ns.list({ tenant: "acme" }); + await ns.table("clicks", { tenant: "acme" }); + await ns.list(); + await ns.table("clicks"); + + const urls = fetchSpy.mock.calls.map((call) => new URL(call[0])); + expect(urls[0].pathname + urls[0].search).toBe("/v1/ops/dlq/stats?tenant=acme"); + expect(urls[1].searchParams.get("table")).toBe("clicks"); + expect(urls[1].searchParams.get("tenant")).toBe("acme"); + expect(urls[2].search).toBe(""); + expect(urls[3].search).toBe("?table=clicks"); + }); + it("stream() delegates to createStream", () => { const ctrl = {} as any; mockStream.mockReturnValue(ctrl); diff --git a/clients/ts/src/types.ts b/clients/ts/src/types.ts index 5feae56d..158d5c44 100644 --- a/clients/ts/src/types.ts +++ b/clients/ts/src/types.ts @@ -471,8 +471,8 @@ export interface PipeRequestOptions { /** * Options for a call to one of the admin routes that address a tenant: * `wh.pipes.list()`, `wh.pipes.get()`, `wh.settings.reload()`, - * `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and - * `wh.sql()`. + * `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()`, + * `wh.dlq.list()`, `wh.dlq.table()` and `wh.sql()`. */ export interface OpsRequestOptions { signal?: AbortSignal; diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 75311090..1634aaab 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -745,14 +745,16 @@ Triggers an immediate re-discovery of the `?tenant=`'s ClickHouse table schemas #### `GET /v1/ops/dlq/stats` — DLQ Statistics -Returns per-table message counts in the Dead Letter Queue — a table's count summed across tenants, since one queue serves every tenant until each has its own. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); the stream and this endpoint always exist. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. +Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, for as long as its queue is kept. The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream exists from the moment the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. **Error responses:** | Status | Body | Cause | | ------ | ---- | ----- | | 401 | `{"error":"invalid token"}` / `{"error":"token expired"}` | A present-but-invalid/expired token was supplied and denied (the gate surfaces the token reason) | +| 400 | `{"error":"invalid query string: …"}` / `{"error":"invalid ?tenant: …"}` | The query string does not parse (`?tenant=acme;x=1`, a bad `%` escape), or `tenant` is empty, repeated, or not a tenant id | | 403 | `{"error":"forbidden"}` | Caller's role is not the policy `admin_role` (`"admin"` by default) | +| 404 | `{"error":"no dead-letter queue for tenant: "}` | The tenant has no dead-letter queue: it has never been served on this data directory, or the id names no tenant | | 500 | `{"error":"stream info failed"}` | NATS JetStream stream-info lookup failed | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while tenant `0`'s JWKS has not been fetched yet (the ops tree verifies as tenant `0`); refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | @@ -760,6 +762,7 @@ Returns per-table message counts in the Dead Letter Queue — a table's count su | Param | Type | Default | Description | | ----- | ---- | ------- | ----------- | +| `tenant` | string | `0` | The tenant whose dead-letter queue is read. | | `table` | string | — | Filter stats to a specific table name (e.g., `?table=clicks` returns only the `clicks` count). | **Response:** @@ -873,9 +876,9 @@ Three values, where the envelope above has four: this is the frame a role restri ## Dead Letter Queue (DLQ) -When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the DLQ NATS stream (`WAVEHOUSE_DLQ`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format` (a pre-v2 message has no `format` field at all, which is how it presents here), or `columns` and `row` that do not pair — is parked without ever reaching a table batch, which is what an operator sees after upgrading across the wire change without draining first. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. +When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the tenant's own DLQ NATS stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format` (a pre-v2 message has no `format` field at all, which is how it presents here), or `columns` and `row` that do not pair — is parked without ever reaching a table batch, which is what an operator sees after upgrading across the wire change without draining first. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. -Use `GET /v1/ops/dlq/stats` to monitor DLQ depth. +Use `GET /v1/ops/dlq/stats` to monitor DLQ depth, per tenant (`?tenant=`). ## Generating a JWT for Testing diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 9d1dc636..6eaf3d54 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -77,20 +77,20 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **router.go** — Route definitions. Public: `/livez`, `/readyz`, and the content-free `/v1/health` SDK ping (plus the permanent `/healthz` alias and the deprecated `/health`, `/ready` aliases). Policy-gated: `/v1/ingest?table={table}`, `/v1/query?table={table}` (structured), `/v1/pipes/{name}` (named pipes), `/v1/stream`. Admin-only (`RequireAdmin` — role == `policy.admin_role`, or a request bearing the operator key's operator bit, which passes even under a nil policy; over a nested settings directory `NewRouter` mounts the gate with no policy at all, whatever `Dependencies.PolicySource` was wired, so the operator key alone passes): `/v1/ops/schema/*`, `/v1/ops/dlq/stats`, `GET /v1/ops/pipes[/{name}]`, `/v1/ops/settings/reload`, `/v1/ops/query` (raw SQL — same gate as the rest of `/v1/ops/*`). - **auth middleware** — the JWT/JWKS authentication middleware is its own package, [`auth/`](#auth--authentication); the router runs it on every `/v1/*` route. -- **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy and the settings reload — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). +- **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. -- **dlq.go** — DLQ stats endpoint (`GET /v1/ops/dlq/stats`): asks `mq.DeadLetterStats.DeadLetterCounts` for the per-table parked counts (optionally one table) and the total. A dead-letter queue that does not exist (`mq.ErrNoDeadLetterQueue`) reads as empty; any other failure to read it is a 500. The queue itself is `internal/mq`'s. +- **dlq.go** — DLQ stats endpoint (`GET /v1/ops/dlq/stats`): asks `mq.DeadLetterStats.DeadLetterCounts` for one tenant's per-table parked counts (optionally one table) and its total — the tenant `?tenant=` names, read strictly by `opsTenant`, tenant `0` without it. The tenant is looked up in the MQ, not the settings registry, so a rejected or removed tenant's parked rows are read like a served one's; a tenant with no dead-letter queue (`mq.ErrNoDeadLetterQueue`) is a 404, and any other failure to read it a 500. The queue itself is `internal/mq`'s. - **health.go** — Liveness (`/livez`), readiness (`/readyz`), and a content-free `Online` ping (`/v1/health`, the SDK's public liveness check); `/healthz` is a permanent alias of `/livez`, and `/health`/`/ready` are deprecated aliases. All three consult an optional `BootState` so they can return 503 while boot-time schema discovery is still failing in the retry loop (see `internal/app`; over a nested directory, while no tenant's has succeeded); once `BootState.Set(nil)` fires, `/livez` returns 200 and stays there. `/readyz` additionally runs a `Ping` each call — `chconn.Pools.Ping` in production: every open pool at once, ready at the first answer, every pool's error joined when none answers; `/v1/health` deliberately does not. ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one resource a process still has one of, the MQ byte budget, follows the default tenant: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`; the zero value when a nested directory has never served a tenant `0`, warned about once at boot), and `onDefaultAdopt` runs its hook only after a reload that adopted it, so another tenant's reload never moves it and a `0` folder that a reload rejects or removes leaves it as it was — like the one setting read per request that follows tenant `0`, the ops gate's admin role. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. Two settings are shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)); and the sweeper keeps the longest `stream.gap_window_minutes` (`longestGapWindow`, read every sweep), since the ingest queue is one stream and a purge is one bound over it — a stream per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 5b) gives each its own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. The `mq.max_bytes_gb` hook only hands the adopted budget to `mq.Broker.SetMaxBytes` under the App's stop context; how it is split across the streams, the time bounds, and the rollback are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -139,16 +139,16 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy. On a bulk-insert failure the batch is re-inserted row by row — except a batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling), which no row could pass and `parkBatch` takes to the DLQ switch whole, logging once per batch rather than twice per row; rows that succeed are acked, and only the rows that fail again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. - **types.go** — `EventMessage` struct (TableName, Scope — reserved, always empty today, ReceivedTimestamp, Format, Columns, Row; `Format` is `FormatJSONCompactEachRow` and `Row` is one positional line whose slots `Columns` names) and `BufferConsumerName` constant, shared across API handlers and the ingest pipeline. - **compact.go** — `EncodeCompactRow`, the positional row encoder every published row goes through, rendering one record over the table's **insertable** columns in declaration order. Serialization only: it validates nothing and judges no value. -- **sweeper.go** — `Sweeper` implements the Active Sweeper pattern. It runs every minute and asks the MQ (`mq.Purger.PurgeAcked`) to drop the ingest events that are **both** ACKed by the buffer consumer (written to ClickHouse) **and** older than the gap window (re-read every sweep: the longest `stream.gap_window_minutes` among the tenants being served — `internal/app`'s `longestGapWindow`). Finding the purge point is `internal/mq`'s (`purge.go`). +- **sweeper.go** — `Sweeper` implements the Active Sweeper pattern. It runs every minute and asks the MQ (`mq.Purger.PurgeAcked`) to drop the ingest events that are **both** ACKed by the buffer consumer (written to ClickHouse) **and** older than the gap window (re-read every sweep: each served tenant's own `stream.gap_window_minutes` — `internal/app`'s `gapWindows` — and none for a tenant no longer served). Finding the purge point is `internal/mq`'s (`purge.go`). ### `mq/` — Message Queue The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. -- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces: `Publisher` (`ErrQueueFull` when the ingest queue is at its byte budget — the API's 503 + `Retry-After`), `Subscriber` (every ingest event, under a named durable consumer — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`; `Consume` delivers on the client goroutine so a blocking handler is backpressure, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer or a closed connection — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic; the caller acks), `DeadLetterStats.DeadLetterCounts` (`ErrNoDeadLetterQueue` when there is none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before a cutoff; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with the byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. -- **subject.go** — The embedded broker's naming, private to the package: the stream names (`WAVEHOUSE`, `WAVEHOUSE_DLQ`), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded. A tail of one token is the form written before the tenant led it and reads as tenant `0`'s table, which is how the events in flight across that upgrade keep inserting, streaming, and counting. -- **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream. Creates stream `WAVEHOUSE` with subjects `ingest.>`, capped at the settings directory's `mq.max_bytes_gb`, and stream `WAVEHOUSE_DLQ` (`dlq.>`, `DiscardOld`) at a tenth of it — always present, since an empty stream costs nothing. `SetMaxBytes` applies a reloaded budget to both live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. Its JetStream calls are bounded to ten seconds, plus five more for the rollback (a budget of its own, not the one that just expired), since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. +- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. It runs on each tenant's stream at that tenant's cutoff. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, or is refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds, plus five more for the rollback (a budget of its own, not the one that just expired), since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index a9c9de1d..8a709426 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -97,7 +97,7 @@ WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so ### Message Queue (NATS) -The stream's disk budget, `mq.max_bytes_gb`, is a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. +Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. **Durability.** The embedded server runs with JetStream `SyncAlways`, so every event is `fsync`'d to disk before `POST /v1/ingest` returns `200`. This makes your storage's `fsync` latency your ingest latency floor — see [Durability & Storage](/durability) to check whether your substrate can sustain it. There is no knob to relax this today ([#139](https://github.com/Wave-RF/WaveHouse/issues/139) tracks a configurable group-commit interval). diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 020c676c..095b4990 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -379,19 +379,19 @@ settings/ └── roles.json ``` -That is the layout a control plane writes. Each folder's `clickhouse` block is its tenant's own ClickHouse, so a tenant answers queries once its first schema discovery against that ClickHouse succeeds (until then its schema-aware routes answer `503`, `schema not loaded yet`); what tenant `0`'s folder still supplies for the whole process — the message queue's budget, the token verifier of the routes that name no tenant, their CORS list — is listed under "What a tenant's folder decides", below. +That is the layout a control plane writes. Each folder's `clickhouse` block is its tenant's own ClickHouse, so a tenant answers queries once its first schema discovery against that ClickHouse succeeds (until then its schema-aware routes answer `503`, `schema not loaded yet`); what tenant `0`'s folder still supplies for the whole process — the token verifier of the routes that name no tenant, their CORS list — is listed under "What a tenant's folder decides", below. The folder name is the tenant id, and each folder is a complete settings directory: everything on the [Settings Directory](/settings-directory) page applies to it as written, except where the rules below say otherwise. The two shapes don't mix — a folder beside the four files, or a loose file beside the folders, is a validation error — and a running server keeps the shape it booted with, so switching is stop, restructure, start. The dedupe store needs no restructuring: it keys every tenant's seen ids by tenant, and the four files are tenant `0`, as a `0` folder is. Dot-prefixed entries are ignored in either shape. `wavehouse validate` checks either shape with the same exit codes; a finding in a nested directory names its folder (`acme/policies.json`), and a folder whose name is not a tenant id is a finding of its own — that folder is skipped, and the rest of the directory still loads. -**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. +**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). Its message queue is kept, at the budget it last had, but the history gap-fill replays is purged from it at the next sweep, as a removed tenant's is, so a stream resumed after the fix has a hole where that history was. The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. -**Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted, so restoring the folder restores the tenant, seen ids included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. +**Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted — its message queue is kept at the budget it last had, and only the history gap-fill replays goes from it, at the next sweep — so restoring the folder restores the tenant, seen ids and parked rows included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. -**The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` reads the queue the whole process shares and ignores the parameter. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). +**The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. The process still has one message queue, and its budget, `mq.max_bytes_gb`, follows tenant `0`'s folder. The queue is shared but addressed per tenant: an event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool too, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. Two settings weigh every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`; and the sweeper, which keeps the longest `stream.gap_window_minutes` among them, since every tenant's events share one message-queue stream and a purge is one bound over it. +**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. -**What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. `mq.max_bytes_gb` stays as tenant `0` last adopted it; tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. The sweeper keeps the longest gap window among the tenants still served, so tenant `0`'s history is purged at theirs, and with no tenant left being served the window is zero, which purges the acknowledged history gap-fill replays. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. +**What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. ### Upgrading behind a proxy that already sends `X-Tenant-ID` @@ -440,13 +440,9 @@ To drain before upgrading: If you skipped the drain, check `wavehouse_ingest_poison_total`, which counts both — `disposition="parked"` is recoverable from `dlq.{table}`, `disposition="dropped"` is gone — see [Dead Letter Queue](#dead-letter-queue-dlq) below. -## Upgrading across the tenant subject token - -Message-queue subjects now lead with the tenant: `ingest.{tenant}.{table}` and `dlq.{tenant}.{table}`, where a settings directory that holds the four files is tenant `0` (`ingest.0.clicks`). The subject change needs no drain of its own; the drain [the envelope upgrade above](#upgrading-across-the-v2-ingest-envelope) asks for still applies. The durable consumers filter `ingest.>`, which the previous subjects match, and a subject with no tenant token reads as tenant `0`'s, so the subject a message arrived on changes nothing about how it is inserted, streamed, or parked, and rows parked under the old `dlq.{table}` keep counting in `GET /v1/ops/dlq/stats`. The one gap is SSE gap-fill, which reads a tenant's own subject: a replay spanning the upgrade omits the events published before it, for one `stream.gap_window_minutes` (15 by default) — the same window as the envelope upgrade's. Clients that need them should backfill over REST. - ## Dead Letter Queue (DLQ) -A failed batch insert is retried row by row; while the tenant's `dlq.enabled` is `true` for the table (the seed default — a hot-reloadable [settings directory](/settings-directory#dead-letter-queue) key, overridable per table), the rows that fail again are published to the `WAVEHOUSE_DLQ` NATS stream under subjects `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) instead of retrying forever. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for, such as by the connection ceiling — skips the row-by-row retry, which no row of it could pass: its tenant's switch is read once for the whole batch, and a tenant no longer served has no switch to read, so its batch is always parked. Monitor DLQ depth via `GET /v1/ops/dlq/stats`. +A failed batch insert is retried row by row; while the tenant's `dlq.enabled` is `true` for the table (the seed default — a hot-reloadable [settings directory](/settings-directory#dead-letter-queue) key, overridable per table), the rows that fail again are published to the tenant's own dead-letter stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) instead of retrying forever. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for, such as by the connection ceiling — skips the row-by-row retry, which no row of it could pass: its tenant's switch is read once for the whole batch, and a tenant no longer served has no switch to read, so its batch is always parked. Monitor DLQ depth via `GET /v1/ops/dlq/stats`, per tenant (`?tenant=`; tenant `0` without it). ## Observability diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 8548ab96..98a8c9c2 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -33,8 +33,8 @@ WaveHouse does not currently expose a knob to relax this — `SyncAlways` is alw Because the publish blocks on `fsync`, **your typical ingest latency is your storage's typical `fsync` latency, and your worst-case publish is your storage's worst-case `fsync`.** When that tail is healthy (sub-millisecond to single-digit milliseconds) the guarantee is essentially free. When it is not, the same code path that handles every production message stalls: - Publishes block for the duration of the `fsync`, so a multi-second `fsync` tail is a multi-second ingest tail. -- The embedded server's stream/consumer setup and every publish run under the JetStream client's request timeout; a slow-enough substrate makes them exceed it. The boot-time symptom is `create stream: ... context deadline exceeded`. -- If the worker cannot drain to ClickHouse faster than producers publish, the stream fills toward [`mq.max_bytes_gb`](/settings-directory#message-queue) and the API returns `503` ([backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs)). +- The embedded server's stream/consumer setup and every publish run under the JetStream client's request timeout; a slow-enough substrate makes them exceed it. The symptom at a first boot, which opens every tenant's queue, is `open dlq stream: ... context deadline exceeded`; a later boot writes nothing, so the first publish is where it shows. +- If the worker cannot drain to ClickHouse faster than producers publish, a tenant's stream fills toward its [`mq.max_bytes_gb`](/settings-directory#message-queue) and the API returns `503` to that tenant ([backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs)). ## Where `SyncAlways` is cheap vs. expensive @@ -100,6 +100,6 @@ If you see any of these, benchmark the `/nats` volume as above: ## See also -- [Settings Directory → Message Queue](/settings-directory#message-queue) — `mq.max_bytes_gb`, the stream's disk budget (hot-reloadable); the SSE gap window inside it is [`stream.gap_window_minutes`](/settings-directory#streaming). +- [Settings Directory → Message Queue](/settings-directory#message-queue) — `mq.max_bytes_gb`, each tenant's queue's disk budget (hot-reloadable); the SSE gap window inside it is [`stream.gap_window_minutes`](/settings-directory#streaming). - [Deployment → Persistent Storage](/deployment#persistent-storage-required-for-containers) — `data_dir` must resolve to a host-backed volume. - [Ingest Pipeline → Backpressure and durability knobs](/ingest-pipeline#backpressure-and-durability-knobs) — the worker-side ack cost and the in-flight backpressure layers. diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index 870133ca..e602e0cf 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -22,15 +22,15 @@ The pipeline is **insert-only**. (Upgrading across the v2 envelope? [Drain the q ## High-level shape -One process consumes a single durable JetStream consumer and fans events out to a goroutine per tenant table — the tenant is the subject's leading token. Each tenant's table batches independently and POSTs to ClickHouse over the HTTP interface (`JSONCompactEachRow`). On a bulk-insert failure the batch is re-inserted row by row, so a single poison row can't sink it: clean rows ack, and only the rows that fail again go to the dead-letter stream. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and meets the dead-letter switch once, whole; a tenant no longer served has no switch to read, so its batch is parked. An envelope the worker cannot *read* — malformed JSON, an unknown row `format` (what a pre-v2 message looks like), or columns and a row that don't pair — never reaches a table loop at all: `parseMsg` parks it on the same dead-letter stream, or, where the DLQ is off for the table, acks and drops it rather than redelivering a message that can never insert. A separate sweeper reclaims stream storage. +Each tenant's events are queued on a JetStream stream of its own. One process holds one durable consumer on each tenant's stream, delivered into one handler, and fans events out to a goroutine per tenant table — the tenant is the subject's leading token. Each tenant's table batches independently and POSTs to ClickHouse over the HTTP interface (`JSONCompactEachRow`). On a bulk-insert failure the batch is re-inserted row by row, so a single poison row can't sink it: clean rows ack, and only the rows that fail again go to the dead-letter stream. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and meets the dead-letter switch once, whole; a tenant no longer served has no switch to read, so its batch is parked. An envelope the worker cannot *read* — malformed JSON, an unknown row `format` (what a pre-v2 message looks like), or columns and a row that don't pair — never reaches a table loop at all: `parseMsg` parks it on the same dead-letter stream, or, where the DLQ is off for the table, acks and drops it rather than redelivering a message that can never insert. A separate sweeper reclaims stream storage. ```mermaid flowchart LR API["POST /v1/ingest"] -->|"publish ingest.TENANT.TABLE"| Stream subgraph NATS["Embedded NATS JetStream (in-process)"] - Stream["WAVEHOUSE stream
all ingest subjects
LimitsPolicy + DiscardNew"] - Cons["buffer-consumer
(durable, pull)"] + Stream["INGEST_TENANT stream, one per tenant
ingest.TENANT.>
LimitsPolicy + DiscardNew"] + Cons["buffer-consumer
(durable, pull, one per tenant stream)"] Stream --> Cons end @@ -46,7 +46,7 @@ flowchart LR TLa -->|"JSONCompactEachRow POST"| CH[("ClickHouse")] TLb --> CH TLc --> CH - TLa -.->|"poison rows"| DLQ["WAVEHOUSE_DLQ
dlq.TENANT.TABLE"] + TLa -.->|"poison rows"| DLQ["DLQ_TENANT stream
dlq.TENANT.TABLE"] D -.->|"unreadable envelope"| DLQ Sweep["Active Sweeper"] -.->|"reads AckFloor, purges"| Stream @@ -204,7 +204,7 @@ Messages still sitting in `msgChan` or the consumer's prefetch buffer at shutdow Delivery can end underneath a running worker: the durable consumer is deleted, or the MQ connection closes. The broker client reports that only through an asynchronous error callback and then stops delivering — no message ever arrives to say so, so a loop that only watches `msgChan` would wait forever while the API kept accepting events nothing writes. `mq.Consumer.Consume` therefore returns a `failed` channel next to `stop` (`mq.ErrDeliveryEnded`, wrapping the broker's reason), and `dispatchLoop` selects on it beside `ctx.Done()` and `msgChan`. On a failure it runs the same bottom-up drain as a shutdown — the rows already in hand are flushed and acked, not abandoned — and then reports the error on the worker's own `failed` channel. A consumer that cannot start at all takes the same path. -The worker does not try to revive the consumer. The app's ingest-worker component returns the error from `app.Run`, which stops every other component and exits non-zero, the same way any failed component does; the supervisor's restart recreates the durable consumer at boot, and everything unacked is redelivered (at-least-once). Passing conditions the client also reports through that callback (a missed heartbeat, a leadership change) are logged at `WARN` and do not end the worker. With the embedded broker (`DontListen`, no external client that could delete the durable) this path is hard to reach today; it matters once a remote broker or per-tenant consumers exist. +The worker does not try to revive the consumer. The app's ingest-worker component returns the error from `app.Run`, which stops every other component and exits non-zero, the same way any failed component does; the supervisor's restart recreates the durable consumer at boot, and everything unacked is redelivered (at-least-once). Passing conditions the client also reports through that callback (a missed heartbeat, a leadership change) are logged at `WARN` and do not end the worker. With the embedded broker (`DontListen`, no external client that could delete a durable) this path is hard to reach; the likeliest way in is a tenant's queue, opened at runtime, that the consumer cannot join. It matters more once a remote broker exists. ## Backpressure and durability knobs @@ -212,23 +212,23 @@ Several layers throttle the pipeline, inner to outer: 1. **`batch`** flushes at `maxBatch` rows or `maxWait`. 2. **`msgChan`** (cap `maxBatch*2`) — when full, the consume callback blocks and delivery pauses. -3. **`pullMaxMessages`** — nats.go's client-side prefetch buffer in front of `msgChan`. -4. **`maxAckPending`** — the server suspends delivery once this many messages are delivered-but-unacked. The outermost in-memory bound. -5. **`MaxBytes` + `DiscardNew`** on the stream (`mq.max_bytes_gb` in the [settings directory](/settings-directory#message-queue), resized in place on reload) — when disk fills (e.g. ClickHouse is down so nothing acks/purges), new publishes are rejected and the API returns 503. +3. **`pullMaxMessages`** — nats.go's client-side prefetch buffer in front of `msgChan`, shared by the tenants' streams (at least one message each). +4. **`maxAckPending`** — the server suspends a tenant's delivery once this many of its messages are delivered-but-unacked; no other tenant's delivery waits on it. The outermost in-memory bound. +5. **`MaxBytes` + `DiscardNew`** on each tenant's stream (its `mq.max_bytes_gb` in the [settings directory](/settings-directory#message-queue), resized in place on reload) — when it fills (e.g. ClickHouse is down so nothing acks/purges), that tenant's new publishes are rejected and the API returns 503. | Knob | Default | Meaning / invariant | | --- | --- | --- | | `maxBatch` | 500 | rows that trigger a flush (soft — coalescing can exceed it) | | `maxWait` | 5s | max time a row waits before its batch flushes | | `ackWait` | 60s | server redelivery timeout; **must exceed `maxWait` + flush time** or in-flight rows get redelivered → duplicate inserts | -| `pullMaxMessages` | 500 | client prefetch; keep `<= maxAckPending` | -| `maxAckPending` | 10,000 | server cap on unacked messages (backpressure) | +| `pullMaxMessages` | 500 | client prefetch, shared by the tenants' streams; keep `<= maxAckPending` | +| `maxAckPending` | 10,000 | server cap on a tenant's unacked messages (backpressure) | `DoubleAck` is used (not fire-and-forget `Ack`) because acking is what records "this data is durably in ClickHouse." With the embedded server's `SyncAlways`, every ack is an fsync and therefore *slow*, which is exactly why acks run in the background (`ackWg`) off the insert path. ## The Active Sweeper -The worker advances the consumer's `AckFloor` by acking; the sweep observes it to decide what is safe to purge. They never call each other — the consumer's `AckFloor` is their only contract. The sweeper (`internal/ingest`) owns the schedule and the window: each tick it calls `mq.Purger.PurgeAcked(buffer-consumer, now − gap window)`, where the window is the longest `stream.gap_window_minutes` among the tenants being served — every tenant's events share one stream and a purge is one bound over it, so purging less is the safe direction until each tenant has its own stream. The steps after the tick below are the embedded broker's implementation of that call. +The worker advances the consumer's `AckFloor` by acking; the sweep observes it to decide what is safe to purge. They never call each other — the consumer's `AckFloor` is their only contract. The sweeper (`internal/ingest`) owns the schedule and the window: each tick it calls `mq.Purger.PurgeAcked(buffer-consumer, cutoffs)` with each served tenant's cutoff at now − its own `stream.gap_window_minutes`; a tenant no longer served — its folder removed or rejected — is given none, and keeps none of the history it has acknowledged. The steps after the tick below are the embedded broker's implementation of that call, run on each tenant's stream at that tenant's cutoff. ```mermaid flowchart TD @@ -236,7 +236,7 @@ flowchart TD Read --> Gap["binary-search the gap-window sequence"] Gap --> Target["target = MIN(ackFloor + 1, gapSeq)"] Target --> Purge["stream.Purge below target"] - Purge -->|"deletes msgs that are BOTH
written to ClickHouse AND past the gap window"| Stream[("WAVEHOUSE stream")] + Purge -->|"deletes msgs that are BOTH
written to ClickHouse AND past the gap window"| Stream[("INGEST_TENANT stream")] ``` `MIN(ackFloor+1, gapSeq)` is the safety argument: never purge past what is in ClickHouse, and never past the SSE replay window. If ClickHouse is down the `AckFloor` stops advancing, purging freezes, and the stream fills toward `MaxBytes` — backpressure by construction. The sweeper is one of `app.Run`'s components (`Sweeper.Start` blocks until the run context is canceled), but an interrupted sweep is harmless and idempotent, so it returns on `ctx.Done()` with no drain of its own — unlike the worker's bounded `stopFunc`. @@ -248,7 +248,7 @@ Today this is a **single-process** design (embedded, in-process NATS — the "co ```mermaid flowchart TD subgraph Cluster["Clustered NATS (Replicas: 3)"] - S["WAVEHOUSE stream"] + S["one shared ingest stream"] end S --> P0["partition 0"] S --> P1["partition 1"] diff --git a/docs/src/content/docs/sdk/admin.md b/docs/src/content/docs/sdk/admin.md index ae1635f7..59db0e5a 100644 --- a/docs/src/content/docs/sdk/admin.md +++ b/docs/src/content/docs/sdk/admin.md @@ -65,6 +65,13 @@ const { data } = await wh.dlq.list(); const { data } = await wh.dlq.table('clicks'); ``` +Each tenant has a dead-letter queue of its own, and the calls read tenant `0`'s without `tenant`. Over [a nested settings directory](/deployment#the-nested-settings-directory), pass `tenant` to read another's — a tenant whose folder was rejected or removed included, since its queue is kept — with the [operator key](/api#authentication), as for the schema reads above. A tenant with no dead-letter queue is a `404`: + +```ts +const { data } = await wh.dlq.list({ tenant: 'acme' }); +const { data: clicks } = await wh.dlq.table('clicks', { tenant: 'acme' }); +``` + `wh.dlq.stream()` exists in the API but is **not yet functional**: there is no server-side DLQ stream today (the SSE bridge only carries `ingest.>` subjects), so it connects and receives no events rather than failing. Live DLQ streaming is tracked in [#197](https://github.com/Wave-RF/WaveHouse/issues/197). --- diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index 56fc09d9..af0cddef 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -104,8 +104,8 @@ createClient(config) → WaveHouseClient ├── .settings (admin) │ └── .reload(opts?) → Promise> ├── .dlq (admin) -│ ├── .list() → Promise> -│ ├── .table(name) → Promise> +│ ├── .list(opts?) → Promise> +│ ├── .table(name, opts?) → Promise> │ └── .stream() → StreamController // not yet functional server-side — #197 └── .sys └── .health() → Promise> diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 9555bf79..c0e7a19b 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -31,7 +31,7 @@ A reload that fails validation is logged (and reported by the endpoint) and the "Previous good settings" is the in-memory snapshot of the running process, nothing more: there is no persisted copy of the files. A restart re-validates the directory from scratch and refuses to start on the same findings the reload rejected, so bad files never survive a restart silently — fix them (or run `wavehouse validate`) before bouncing the server. -A directory that holds one folder per tenant instead of the four files is [a nested settings directory](/deployment#the-nested-settings-directory): each folder is everything this page describes, but it is not watched, a rejected folder stops its tenant being served rather than keeping the previous settings, and the keys the whole process shares are not read from the tenant's own folder: they come from tenant `0`'s, bar the two that weigh every tenant being served: the SSE keepalive, which follows the shortest `stream.keepalive_interval` among them, and the sweeper's gap window, the longest `stream.gap_window_minutes` among them — that section lists which keys. +A directory that holds one folder per tenant instead of the four files is [a nested settings directory](/deployment#the-nested-settings-directory): each folder is everything this page describes, but it is not watched, a rejected folder stops its tenant being served rather than keeping the previous settings, and the keys the whole process shares are not read from the tenant's own folder: they come from tenant `0`'s, bar the one that weighs every tenant being served: the SSE keepalive, which follows the shortest `stream.keepalive_interval` among them — that section lists which keys. Every adoption — boot and every reload — goes through the same `Validate`, so the policy, the roles, and the pipes are checked with the current rules each time they are read; there is no stored copy that can skip validation. All four files are adopted as one snapshot: a request is evaluated against the policy, pipes, and tunables of a single adoption, never a mix. @@ -124,7 +124,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `dedupe.id_field` | `event_id` | Dedup key field — see [Deduplication](#deduplication). | | `dedupe.require_id` | `false` | Reject rows missing the id field — see [Deduplication](#deduplication). | | `dedupe.tables.
.{id_field, require_id}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | -| `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the `WAVEHOUSE_DLQ` stream (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | +| `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the tenant's dead-letter stream (`DLQ_{tenant}`) (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | | `dlq.tables.
.enabled` | `{}` | Optional per-table override of the switch. | | `query.timestamp_bucket_seconds` | `60` | Bucket (seconds, `>= 0`) that a structured query's relative time range is truncated to, so near-identical queries share a cache entry; `0` disables bucketing. Read per query. | | `query.default_max_rows` | `10000` | Fallback result `LIMIT` (`>= 1`) applied to a structured query when the caller and policy specify none. A result-**shaping** default, not a resource limit — server-wide limits (memory, rows scanned, execution time) belong in ClickHouse, see [Server-side resource limits](/configuration#server-side-resource-limits). | @@ -132,7 +132,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `stream.keepalive_interval` | `30` | Seconds (`>= 1`) a quiet `GET /v1/stream` connection may go without a write before the server sends a `:` keepalive comment — keep it under your proxy's idle timeout; see [Streaming](#streaming). | | `stream.keepalive_buckets` | `3` | Load-spreading (`>= 1`): connections are spread across N buckets so each tick nudges ~1/N of live streams. Most deployments leave it. | | `stream.gap_window_minutes` | `15` | Minutes (`>= 0`) of written-to-ClickHouse history the Active Sweeper keeps in NATS for `Last-Event-ID` gap-fill; applies from the next sweep. | -| `mq.max_bytes_gb` | `50` | Disk budget (GB, `>= 1`) for the embedded NATS `WAVEHOUSE` ingest stream; the `WAVEHOUSE_DLQ` stream gets a tenth of it. A reload updates the live streams in place. See [Message Queue](#message-queue). | +| `mq.max_bytes_gb` | `50` | Disk budget (GB, `>= 1`) for the tenant's embedded NATS ingest stream (`INGEST_{tenant}`); its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. A reload updates the live streams in place. See [Message Queue](#message-queue). | | `cors.allowed_origins` | `["*"]` | Allowed CORS origins, applied per request. `"*"` allows any browser origin. WaveHouse is a Bearer-token API — `Access-Control-Allow-Credentials` is intentionally never sent, so this allowlist controls *which origins can read responses*, not cookie scope. Tighten to your frontend's exact origin(s) in production (e.g. `["https://dashboard.example.com", "http://localhost:3000"]`). An empty list `[]` denies every browser origin (no `Access-Control-Allow-Origin` is ever sent); `"*"` is the only allow-all spelling. Over [a nested settings directory](/deployment#the-nested-settings-directory) each tenant's list decorates its own responses, the preflight included; which list answers a preflight, the tenant-exempt routes, and a refused request is [spelled out there](/deployment#multi-tenant-deployments). | ```json @@ -212,16 +212,16 @@ The `auth` block is the verifier wiring, minus the secrets. `jwks_url` (absolute A failed batch insert is retried row by row; a row that fails again on its own is a poison row. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the retry, which no row of it could pass, and every row of it is a poison row. `dlq.enabled` (seed default `true`) decides what happens to it, resolved per table (`dlq.tables.
.enabled` → global) at the moment of the failure, so a reload applies to the next poison row: -- `true` — the row is published to the `WAVEHOUSE_DLQ` NATS stream under `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) with the failure in its headers, and its original is acked. Inspect it with `GET /v1/ops/dlq/stats` (admin-only). +- `true` — the row is published to the tenant's dead-letter stream (`DLQ_{tenant}`) under `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) with the failure in its headers, and its original is acked. Inspect it with `GET /v1/ops/dlq/stats` (admin-only; `?tenant=` names the tenant). - `false` — the row is left unacked, so NATS redelivers it and it retries until it inserts or the switch is flipped back. For every row the worker **can read**, nothing is ever dropped either way — the choice is *park it* versus *keep retrying*. **One exception, new in this release:** an envelope the worker cannot read *at all* — malformed JSON, an unknown `format` (what a pre-v2 in-flight message looks like), or `columns` and `row` that do not pair — can never insert, so redelivering it forever would wedge the consumer. With the DLQ off for the table it is acked and **dropped**, logged at `ERROR` and counted by `wavehouse_ingest_poison_total` with `disposition="dropped"` (also labeled by `table` and `reason`; an envelope parked on the DLQ carries `disposition="parked"`). See [Ingest Pipeline](/ingest-pipeline) — and drain the ingest queue before upgrading. -For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in the shared ingest queue, where it would stop the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging it. +For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in its ingest queue, redelivered for as long as the tenant is away and stopping the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging that queue. -The `WAVEHOUSE_DLQ` stream always exists (an empty stream costs nothing) and the stats endpoint is always registered — the switch is purely behavioral, which is what makes it safe to reload. +A tenant's dead-letter stream exists from the moment the tenant is first served (an empty stream costs nothing) and the stats endpoint is always registered — the switch is purely behavioral, which is what makes it safe to reload. ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the embedded JetStream `WAVEHOUSE` stream that buffers ingested events until the worker writes them to ClickHouse; the `WAVEHOUSE_DLQ` stream gets a tenth of it. The stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the stream refuse new publishes until the worker drains it back under the limit — nothing already accepted is dropped. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Size it from [Durability & Storage](/durability). +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the worker drains it back under the limit — nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to — so a queue fills until its budget or the disk runs out, whichever comes first; size them together from [Durability & Storage](/durability). ## Streaming @@ -229,4 +229,4 @@ The `WAVEHOUSE_DLQ` stream always exists (an empty stream costs nothing) and the - `stream.keepalive_interval` (seed default `30`) — seconds a quiet connection may go without a write before the server sends a `:` keepalive comment. It exists to stay under whatever idle timeout sits between WaveHouse and the client; the default clears the common 55–60s proxy windows with margin, and a tighter edge (Azure Application Gateway 20s, CloudFront 30s) wants a lower value — see [Behind a reverse proxy → Idle timeouts](/reverse-proxy#idle-timeouts-by-provider). A reload rebuilds the keepalive wheel in place: live connections stay open and are redistributed across the new ring, each getting at most one full new period before its next keepalive. - `stream.keepalive_buckets` (seed default `3`) — spreads the keepalive writes across the interval (one bucket fires every `keepalive_interval ÷ keepalive_buckets`) so the server nudges ~1/N of connections per tick instead of all at once. It changes only how the writes are spread in time, never the period. -- `stream.gap_window_minutes` (seed default `15`) — minutes of already-written-to-ClickHouse history the Active Sweeper keeps in NATS so a reconnecting client's `Last-Event-ID` replay can bridge the gap; a drop longer than this resumes with a hole. Bounded by the stream's disk budget, [`mq.max_bytes_gb`](#message-queue). A reload applies from the next sweep (every minute). +- `stream.gap_window_minutes` (seed default `15`) — minutes of already-written-to-ClickHouse history the Active Sweeper keeps in NATS so a reconnecting client's `Last-Event-ID` replay can bridge the gap; a drop longer than this resumes with a hole. Bounded by the tenant's disk budget, [`mq.max_bytes_gb`](#message-queue). A reload applies from the next sweep (every minute). diff --git a/docs/src/content/docs/why-wavehouse.md b/docs/src/content/docs/why-wavehouse.md index a5b63e21..26ac9d70 100644 --- a/docs/src/content/docs/why-wavehouse.md +++ b/docs/src/content/docs/why-wavehouse.md @@ -53,7 +53,7 @@ Even if you remember to batch client-side, a naive ingest path has no safe way t - **No backpressure channel.** If the merger falls behind, ClickHouse raises an error at the *next* insert. The client has already left. - **No DLQ.** Bad events that fail to insert are either lost or logged into ClickHouse's error log. Good luck replaying yesterday's dropped rows. -WaveHouse fixes all three at the gateway: validates every payload against the real `system.columns` schema before accepting, returns `503 Service Unavailable` with a `Retry-After` header when the NATS WAL fills, and routes failed batch inserts to a dedicated `WAVEHOUSE_DLQ` stream you can inspect via `GET /v1/ops/dlq/stats`. +WaveHouse fixes all three at the gateway: validates every payload against the real `system.columns` schema before accepting, returns `503 Service Unavailable` with a `Retry-After` header when the NATS WAL fills, and routes failed batch inserts to a dedicated dead-letter stream, one per tenant, you can inspect via `GET /v1/ops/dlq/stats`. ### No real-time push @@ -156,7 +156,7 @@ flowchart TB | Real-time push | WebSocket service + bridge from Kafka | Built in (`/v1/stream`) | | Schema validation | Custom code in ingest API | Built in (discovers `system.columns`) | | Row/column access control | Custom middleware or a dedicated service | Built in (Hasura-style, JWT-driven) | -| Dead letter queue | Custom retry + dead topic on Kafka | Built in (`WAVEHOUSE_DLQ`) | +| Dead letter queue | Custom retry + dead topic on Kafka | Built in (a dead-letter stream per tenant) | | Client SDK | Each team writes one | `@wavehouse/sdk` (TypeScript, one dependency, codegen) | The DIY path works — big teams run it — but the ops cost is not small. You're paying for a Kafka cluster (or Confluent bill), a second service you wrote from scratch, and all the debugging hours when the batching consumer stalls at 3 a.m. @@ -190,7 +190,7 @@ Tinybird wins on "zero ops to start." WaveHouse wins on "own your data plane and | Self-hosted | ✓ | ✓ | ✗ | ✓ | | Handles N-row inserts safely | ✗ merge blowup | ✓ via Kafka | ✓ | ✓ native | | Schema validation at the edge | ✗ | Custom | ✓ | ✓ (discovers schema) | -| Dead letter queue | ✗ | Custom | Partial | ✓ `WAVEHOUSE_DLQ` | +| Dead letter queue | ✗ | Custom | Partial | ✓ dead-letter stream per tenant | | Backpressure (503 + Retry-After) | ✗ | Custom | ✓ | ✓ | | Idempotent ingest (dedup by ID) | ✗ | Custom | ✓ | ✓ optional | | Real-time push (SSE) | ✗ | Custom service | ✗ | ✓ native, gap-fill | @@ -221,7 +221,7 @@ flowchart TB NATS --> BC["Buffer consumer
5-second batches"]:::wh BC --> CH[("ClickHouse")]:::store - BC -. "on failure" .-> DLQ["WAVEHOUSE_DLQ"]:::fail + BC -. "on failure" .-> DLQ["dead-letter stream"]:::fail ``` **Query path with tiered cache:** diff --git a/internal/api/dlq.go b/internal/api/dlq.go index 9de69ad5..267d2dc4 100644 --- a/internal/api/dlq.go +++ b/internal/api/dlq.go @@ -7,6 +7,7 @@ import ( "net/http" "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/tenant" ) // DLQHandler exposes Dead Letter Queue statistics. @@ -18,18 +19,30 @@ func NewDLQHandler(stats mq.DeadLetterStats) *DLQHandler { return &DLQHandler{Counts: stats} } -// Stats returns per-table message counts on the dead-letter queue. -// Supports optional ?table= query parameter to filter by table name. +// Stats returns per-table message counts on one tenant's dead-letter queue: +// the tenant ?tenant= names (read strictly, as every ops read does — +// opsTenant), tenant.Default without it. The queue is the MQ's, not the +// settings', so it is looked up there: a tenant whose folder was rejected or +// removed is read like one being served, for as long as its queue is kept, +// and an id with no queue is a 404. Supports optional ?table= query parameter +// to filter by table name. func (h *DLQHandler) Stats(w http.ResponseWriter, r *http.Request) { - counts, err := h.Counts.DeadLetterCounts(r.Context(), r.URL.Query().Get("table")) + id, named, ok := opsTenant(w, r) + if !ok { + return + } + if !named { + id = tenant.Default + } + counts, err := h.Counts.DeadLetterCounts(r.Context(), id, r.URL.Query().Get("table")) if err != nil { - if !errors.Is(err, mq.ErrNoDeadLetterQueue) { - slog.ErrorContext(r.Context(), "dlq stats failed", "error", err) - writeJSONError(w, http.StatusInternalServerError, "stream info failed") + if errors.Is(err, mq.ErrNoDeadLetterQueue) { + writeJSONError(w, http.StatusNotFound, "no dead-letter queue for tenant: "+id.String()) return } - // No dead-letter queue: nothing can have been parked. - counts = mq.DeadLetterCounts{Tables: map[string]uint64{}} + slog.ErrorContext(r.Context(), "dlq stats failed", "tenant", id, "error", err) + writeJSONError(w, http.StatusInternalServerError, "stream info failed") + return } w.Header().Set("Content-Type", "application/json") diff --git a/internal/api/dlq_test.go b/internal/api/dlq_test.go index 4023e354..d814cd8b 100644 --- a/internal/api/dlq_test.go +++ b/internal/api/dlq_test.go @@ -15,56 +15,53 @@ import ( "github.com/stretchr/testify/require" ) -// parkedMsg is a message as the ingest worker would hand it to the DLQ. -func parkedMsg(table string) *mq.Message { +// parkedMsg is a message as the ingest worker would hand it to the DLQ, +// parked under tenant id's table. +func parkedMsg(id tenant.ID, table string) *mq.Message { return (&testutil.MockMessage{ - MsgTopic: mq.Topic{Tenant: tenant.Default, Table: table}, + MsgTopic: mq.Topic{Tenant: id, Table: table}, MsgData: []byte(`{"table_name":"` + table + `"}`), }).Message() } -func TestDLQStats_EmptyWhenNoStream(t *testing.T) { - // The embedded MQ always has a dead-letter queue, so its absence comes - // from a mock. - handler := NewDLQHandler(&testutil.MockDeadLetterStats{Err: mq.ErrNoDeadLetterQueue}) - - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) +// dlqStats serves GET /v1/ops/dlq/stats with query through handler. +func dlqStats(t *testing.T, handler *DLQHandler, query string) *httptest.ResponseRecorder { + t.Helper() + req := httptest.NewRequestWithContext(t.Context(), http.MethodGet, "/v1/ops/dlq/stats"+query, nil) rec := httptest.NewRecorder() - handler.Stats(rec, req) + return rec +} - assert.Equal(t, http.StatusOK, rec.Code) - - var resp map[string]any - require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &resp)) - - tables, ok := resp["tables"].(map[string]any) - require.True(t, ok) - assert.Empty(t, tables) - assert.Equal(t, float64(0), resp["total"]) +// A tenant with no dead-letter queue — never given one on this data +// directory, or an id nobody uses — is a 404 that names it, not an empty +// count that would read as "nothing parked" for a typo. +func TestDLQStats_NoQueueIs404(t *testing.T) { + // The embedded MQ opens a served tenant's queue at boot, so a queue's + // absence comes from a mock. + stats := &testutil.MockDeadLetterStats{Err: mq.ErrNoDeadLetterQueue} + + rec := dlqStats(t, NewDLQHandler(stats), "?tenant=acmee") + assert.Equal(t, http.StatusNotFound, rec.Code) + assert.Contains(t, rec.Body.String(), "no dead-letter queue for tenant: acmee") + testutil.AssertJSONErrorResponse(t, rec) + assert.Equal(t, tenant.ID("acmee"), stats.Tenant) } func TestDLQStats_ReturnsCorrectCounts(t *testing.T) { - dir := t.TempDir() - emb, err := mq.NewEmbedded(dir, 1024*1024) - require.NoError(t, err) - defer func() { _ = emb.Close() }() + emb := testutil.NewEmbeddedMQ(t, 1024*1024) ctx := context.Background() // Park messages on the dead-letter queue. for i := 0; i < 3; i++ { - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("events"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "events"))) } for i := 0; i < 2; i++ { - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("users"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "users"))) } - handler := NewDLQHandler(emb) - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) + rec := dlqStats(t, NewDLQHandler(emb), "") assert.Equal(t, http.StatusOK, rec.Code) @@ -78,21 +75,20 @@ func TestDLQStats_ReturnsCorrectCounts(t *testing.T) { assert.Equal(t, float64(5), resp["total"]) } +func TestDLQStats_EmptyBeforeAnyFailure(t *testing.T) { + rec := dlqStats(t, NewDLQHandler(testutil.NewEmbeddedMQ(t, 1024*1024)), "") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{},"total":0}`, rec.Body.String()) +} + func TestDLQStats_SingleTable(t *testing.T) { - dir := t.TempDir() - emb, err := mq.NewEmbedded(dir, 1024*1024) - require.NoError(t, err) - defer func() { _ = emb.Close() }() + emb := testutil.NewEmbeddedMQ(t, 1024*1024) ctx := context.Background() - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("orders"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "orders"))) - handler := NewDLQHandler(emb) - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) + rec := dlqStats(t, NewDLQHandler(emb), "") assert.Equal(t, http.StatusOK, rec.Code) @@ -105,33 +101,61 @@ func TestDLQStats_SingleTable(t *testing.T) { } func TestDLQStats_BrokerFailureIsAnError(t *testing.T) { - handler := NewDLQHandler(&testutil.MockDeadLetterStats{Err: errors.New("broker unavailable")}) - - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) - + rec := dlqStats(t, NewDLQHandler(&testutil.MockDeadLetterStats{Err: errors.New("broker unavailable")}), "") assert.Equal(t, http.StatusInternalServerError, rec.Code, "a failed read is not an empty queue") } func TestDLQStats_PassesTheTableFilter(t *testing.T) { - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - defer func() { _ = emb.Close() }() + emb := testutil.NewEmbeddedMQ(t, 1024*1024) ctx := context.Background() - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("default.orders"))) - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("users"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "default.orders"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "users"))) - handler := NewDLQHandler(emb) - req := httptest.NewRequestWithContext(ctx, http.MethodGet, "/v1/ops/dlq/stats?table=default.orders", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) + rec := dlqStats(t, NewDLQHandler(emb), "?table=default.orders") var resp map[string]any require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &resp)) assert.Equal(t, map[string]any{"default.orders": float64(1)}, resp["tables"]) assert.Equal(t, float64(2), resp["total"]) } + +// ?tenant= reads that tenant's queue alone, and no parameter reads tenant 0's +// — the ops-read convention. The handler asks the MQ, not the settings, so a +// tenant the settings no longer serve (here, none at all) is read by name +// for as long as its queue is kept. +func TestDLQStats_ReadsTheNamedTenantsQueue(t *testing.T) { + emb := testutil.NewEmbeddedMQ(t, 1024*1024, tenant.Default, "acme") + ctx := context.Background() + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "events"))) + for range 2 { + require.NoError(t, emb.DeadLetter(ctx, parkedMsg("acme", "events"))) + } + handler := NewDLQHandler(emb) + + rec := dlqStats(t, handler, "?tenant=acme") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{"events":2},"total":2}`, rec.Body.String()) + + rec = dlqStats(t, handler, "?tenant=acme&table=users") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{},"total":2}`, rec.Body.String()) + + rec = dlqStats(t, handler, "") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{"events":1},"total":1}`, rec.Body.String(), "no parameter reads tenant 0") + + rec = dlqStats(t, handler, "?tenant=globex") + assert.Equal(t, http.StatusNotFound, rec.Code, "a tenant with no queue") +} + +// The parameter is read strictly, like every ops read's (opsTenant): a +// query that misparses must not fall back to tenant 0's counts. +func TestDLQStats_RefusesAMalformedTenant(t *testing.T) { + for _, query := range []string{"?tenant=a.b", "?tenant=", "?tenant=a&tenant=b", "?tenant=acme;x=1"} { + stats := &testutil.MockDeadLetterStats{} + rec := dlqStats(t, NewDLQHandler(stats), query) + assert.Equal(t, http.StatusBadRequest, rec.Code, query) + assert.Empty(t, stats.Tenant, "%s: nothing is read", query) + } +} diff --git a/internal/api/ingest.go b/internal/api/ingest.go index c429aaa4..029799b6 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -707,7 +707,7 @@ func (h *IngestHandler) processRecord( slog.DebugContext(ctx, "publishing event to the ingest queue", "table", table, "scope", scope) if err := h.Publisher.Publish(ctx, mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope}, payload); err != nil { if errors.Is(err, mq.ErrQueueFull) { - slog.WarnContext(ctx, "ingest queue is full", "table", table, "scope", scope) + slog.WarnContext(ctx, "ingest queue is full", "error", err, "table", table, "scope", scope) return false, nil, &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} } slog.ErrorContext(ctx, "failed to publish to the ingest queue", "error", err, "table", table, "scope", scope) diff --git a/internal/api/router_test.go b/internal/api/router_test.go index e0240ba6..03a39c76 100644 --- a/internal/api/router_test.go +++ b/internal/api/router_test.go @@ -15,7 +15,6 @@ import ( "github.com/Wave-RF/WaveHouse/internal/auth" "github.com/Wave-RF/WaveHouse/internal/discovery" - "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" "github.com/Wave-RF/WaveHouse/internal/settings" @@ -332,9 +331,7 @@ func TestNewRouter_RoutesRegistered(t *testing.T) { pub := &testutil.MockPublisher{} hub := stream.NewHub(nil, nil, nil) - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) deps := Dependencies{ Tenants: testTenants(), diff --git a/internal/app/app.go b/internal/app/app.go index f853a76b..dc15bf5d 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -18,8 +18,7 @@ // handed whole to each component's wiring function, which derives the // per-call getters the internal packages take: keyed by the request's store // for the handlers, by tenant id for the async paths (perTenant), and fixed -// to the default tenant for the process-wide resources #583 has not yet made -// per tenant (defaultSetting). +// to the default tenant for the ops gate of a flat directory (defaultSetting). package app import ( @@ -90,9 +89,9 @@ type App struct { listener net.Listener // tenants is the registry every tenant-aware path resolves through, and - // the owner of every reload. The one process-wide resource left, the MQ, - // still follows its default tenant, through defaultStore: tenant 0's - // store as of its last adoption (defaultSetting). + // the owner of every reload. defaultStore is tenant 0's store as of its + // last adoption, which the ops gate of a flat directory reads its admin + // role from (defaultSetting). tenants *settings.Registry defaultStore atomic.Pointer[settings.Store] // policies is the default tenant's policy, for the ops gate of a flat diff --git a/internal/app/app_test.go b/internal/app/app_test.go index fc5d2755..2ecd77d9 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -237,7 +237,7 @@ func TestReload_DrivesTheRegisteredHooks(t *testing.T) { a := newApp(t, cfg, Options{}) dedup := a.dedup.For(tenant.Default) require.False(t, dedup.Open()) - require.Equal(t, int64(1<<30), a.mq.MaxBytes()) + require.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, dir, map[string]any{ "dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}, @@ -246,8 +246,8 @@ func TestReload_DrivesTheRegisteredHooks(t *testing.T) { _, adopted := a.tenants.Reload("test") require.True(t, adopted) assert.True(t, dedup.Open(), "dedupe hook opened the store") - // How the budget is split across the MQ's queues is internal/mq's to test. - assert.Equal(t, int64(2<<30), a.mq.MaxBytes(), "mq hook applied the new byte budget") + // How the budget is split across the tenant's queues is internal/mq's to test. + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default), "mq hook applied the new byte budget") rewriteSettings(t, dir, map[string]any{"mq": map[string]any{"max_bytes_gb": 2}}) _, adopted = a.tenants.Reload("test") @@ -397,12 +397,13 @@ func TestNew_NestedWithoutAnOperatorKeyWarnsTheOpsTreeIsClosed(t *testing.T) { }) } -// The process-wide resources follow tenant 0 alone: another tenant's reload -// never moves them, and a rejected 0 folder leaves them as they were rather -// than reconfiguring them from nothing. The dedupe stores are per tenant -// (story 7), so each follows its own folder instead — the contrast the -// same reloads show. -func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { +// A tenant's queue budget and dedupe store follow its own folder alone: +// another tenant's reload moves neither. A rejected or removed folder keeps +// its tenant's queue at the budget it last had — removing never touches +// data — while its dedupe store closes, its seen ids kept. CORS is read per +// request, so a lost 0 folder is felt at once on the routes that read tenant +// 0's list. +func TestReload_NestedHooksFollowEachTenant(t *testing.T) { dedupeOn := map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}} grown := map[string]any{"dedupe": dedupeOn, "mq": map[string]any{"max_bytes_gb": 2}} root := writeNestedSettings(t, map[string]map[string]any{ @@ -413,7 +414,8 @@ func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { dedup0, dedupAcme := a.dedup.For(tenant.Default), a.dedup.For("acme") require.False(t, dedup0.Open()) require.False(t, dedupAcme.Open()) - require.Equal(t, int64(1<<30), a.mq.MaxBytes()) + require.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) + require.Equal(t, int64(1<<30), a.mq.MaxBytes("acme"), "each served tenant's queue opens at boot at its own budget") // CORS is per tenant, not a hook's: a tenant route reads its own tenant's // list and the exempt routes tenant 0's (the seed's ["*"] in every folder // here), both through the registry, so a lost 0 folder is felt at once. @@ -434,26 +436,27 @@ func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { _, adopted := a.tenants.Reload("test") require.True(t, adopted) assert.True(t, dedupAcme.Open(), "acme's dedupe switch opens acme's own store") + assert.Equal(t, int64(2<<30), a.mq.MaxBytes("acme"), "acme's budget resizes acme's own queue") assert.False(t, dedup0.Open(), "and moves nothing of tenant 0's") - assert.Equal(t, int64(1<<30), a.mq.MaxBytes()) + assert.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, filepath.Join(root, "0"), grown) _, adopted = a.tenants.Reload("test") require.True(t, adopted) assert.True(t, dedup0.Open()) - assert.Equal(t, int64(2<<30), a.mq.MaxBytes()) + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, filepath.Join(root, "0"), invalidQuery) _, adopted = a.tenants.Reload("test") require.False(t, adopted) assert.False(t, dedup0.Open(), "a rejected 0 folder closes tenant 0's own store, which answers no request now") assert.True(t, dedupAcme.Open(), "and costs acme nothing") - assert.Equal(t, int64(2<<30), a.mq.MaxBytes(), "the process-wide budget stays as tenant 0 last adopted it") + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default), "tenant 0's queue is kept at the budget it last had") assert.Empty(t, allowOrigin("/version"), "the exempt routes read tenant 0 through the registry, which is no longer serving it") assert.Equal(t, "*", allowOrigin("/v1/health", "acme"), "acme's own routes keep acme's list") - // A removed 0 folder is the same: the registry forgets the tenant, the - // process keeps the wiring it last adopted. + // A removed 0 folder is the same: the registry forgets the tenant, and + // its queue stays at the budget it last had. require.NoError(t, os.RemoveAll(filepath.Join(root, "0"))) _, adopted = a.tenants.Reload("test") require.True(t, adopted) @@ -461,7 +464,8 @@ func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { require.False(t, known) assert.False(t, dedup0.Open()) assert.True(t, dedupAcme.Open()) - assert.Equal(t, int64(2<<30), a.mq.MaxBytes()) + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default)) + assert.Equal(t, int64(2<<30), a.mq.MaxBytes("acme")) assert.Empty(t, allowOrigin("/version")) assert.Equal(t, "*", allowOrigin("/v1/health", "acme")) } @@ -682,12 +686,11 @@ func gapWindow(minutes int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": minutes}} } -// One ingest stream holds every tenant's events and the sweeper purges below -// one sequence, so it keeps the longest gap window among the tenants being -// served — every tenant's gap-fill history is inside it (a stream per tenant -// will honor each tenant's own, #583 story 5b). A flat directory's single +// Each tenant being served keeps its own stream.gap_window_minutes, since +// each has a queue of its own; a rejected tenant is not served, so it is not +// named and keeps no history (mq.Purger.PurgeAcked). A flat directory's single // tenant gets exactly its own window. -func TestLongestGapWindow(t *testing.T) { +func TestGapWindows(t *testing.T) { open := func(t *testing.T, dir string) *settings.Registry { t.Helper() guardGlobals(t) @@ -697,22 +700,21 @@ func TestLongestGapWindow(t *testing.T) { } t.Run("flat directory", func(t *testing.T) { - assert.Equal(t, 45*time.Minute, longestGapWindow(open(t, writeSettings(t, gapWindow(45))))) + assert.Equal(t, map[tenant.ID]time.Duration{tenant.Default: 45 * time.Minute}, gapWindows(open(t, writeSettings(t, gapWindow(45))))) }) t.Run("nested directory", func(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": gapWindow(15), "globex": gapWindow(60), "initech": gapWindow(30)}) tenants := open(t, root) - assert.Equal(t, 60*time.Minute, longestGapWindow(tenants)) + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "globex": 60 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants)) - // A rejected tenant is not being served, so its window is not weighed. rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) tenants.Reload("test") - assert.Equal(t, 30*time.Minute, longestGapWindow(tenants)) + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants)) }) - t.Run("no tenant served keeps nothing", func(t *testing.T) { - assert.Zero(t, longestGapWindow(open(t, writeNestedSettings(t, map[string]map[string]any{"acme": invalidQuery})))) + t.Run("no tenant served names none", func(t *testing.T) { + assert.Empty(t, gapWindows(open(t, writeNestedSettings(t, map[string]map[string]any{"acme": invalidQuery})))) }) } diff --git a/internal/app/wire.go b/internal/app/wire.go index 009f9920..60cbdcd8 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -65,18 +65,13 @@ func (a *App) wireSettings() error { return fmt.Errorf("settings directory %s invalid, refusing to start — findings above; `wavehouse validate` reproduces them, `wavehouse bootstrap` writes a starter directory", a.cfg.Settings.Dir) } a.tenants = tenants - // Registered first: hooks run in registration order, and every other one - // reads tenant 0 through the store this one tracks. + // Registered first: hooks run in registration order, so every reload + // updates the tracked store before any other hook runs. a.trackDefaultStore() a.onDefaultAdopt(a.trackDefaultStore) a.policies = func() *policy.Policy { return defaultSetting(a, (*settings.Store).Policy) } - switch _, served := tenants.For(tenant.Default); { - case !tenants.Nested(): - if a.policies() == nil { - slog.Warn("no policy adopted — every token-based request is denied until policies.json defines one (fail closed)") - } - case !served: - slog.Warn("nested settings directory with no tenant 0 being served: the MQ byte budget is still configured from tenant 0's config.json, so it runs unconfigured until a 0 folder is adopted") + if !tenants.Nested() && a.policies() == nil { + slog.Warn("no policy adopted — every token-based request is denied until policies.json defines one (fail closed)") } return nil } @@ -84,20 +79,18 @@ func (a *App) wireSettings() error { // trackDefaultStore remembers tenant 0's store as of its last adoption. The // registry stops handing out a rejected tenant's store and forgets a removed // one, but the store keeps its last adopted document either way — and that is -// what the process-wide resources go on following (defaultSetting). +// what defaultSetting goes on reading. func (a *App) trackDefaultStore() { if store, ok := a.tenants.For(tenant.Default); ok { a.defaultStore.Store(store) } } -// defaultSetting reads one setting of the default tenant, which the one -// process-wide resource left (the MQ) follows until #583 gives each tenant -// its own. It reads tenant 0's last adopted document, so a -// 0 folder a reload rejected or removed leaves every reader as it was — -// the MQ's byte budget a hook reconciles and the one read per request (the ops -// gate's admin role) alike. A nested directory that has never served a tenant -// 0 reads T's zero value, which wireSettings warned about at boot. +// defaultSetting reads one setting of the default tenant: the admin role the +// ops gate of a flat directory reads per request. It reads tenant 0's last +// adopted document, so a 0 folder a reload rejected or removed leaves its +// reader as it was. A nested directory that has never served a tenant 0 reads +// T's zero value, and its ops gate reads no policy at all. func defaultSetting[T any](a *App, get func(*settings.Store) T) T { store := a.defaultStore.Load() if store == nil { @@ -108,8 +101,8 @@ func defaultSetting[T any](a *App, get func(*settings.Store) T) T { } // onDefaultAdopt registers fn to run after each reload that adopts the -// default tenant, so a nested directory's other tenants never move the -// process-wide resources, and a rejected 0 folder leaves them as they were. +// default tenant, so a nested directory's other tenants never move what +// follows it, and a rejected 0 folder leaves that as it was. func (a *App) onDefaultAdopt(fn func()) { a.tenants.AfterAdopt(func(adopted []tenant.ID) { if slices.Contains(adopted, tenant.Default) { @@ -137,20 +130,16 @@ func shortestKeepalive(tenants *settings.Registry) (period time.Duration, bucket return period, buckets } -// longestGapWindow is the shape of the one purge bound every tenant's events -// share: the ingest queue is one stream and the sweeper purges below one -// sequence, so the history kept is the longest stream.gap_window_minutes -// among the tenants being served — purging less, never more, so every -// tenant's gap-fill history survives — at the cost of one tenant holding the -// others' history for longer, which a stream per tenant will end (#583 story -// 5b). A flat directory's one tenant gets exactly its own window; -// with no tenant served the zero window purges everything acknowledged. -func longestGapWindow(tenants *settings.Registry) time.Duration { - var window time.Duration - for _, store := range tenants.All() { - window = max(window, store.GapWindow()) +// gapWindows is the history the sweeper keeps for each tenant being served: +// its own stream.gap_window_minutes, since each tenant's events have a queue +// of their own. A tenant it does not name — removed or rejected — keeps no +// history (mq.Purger.PurgeAcked). +func gapWindows(tenants *settings.Registry) map[tenant.ID]time.Duration { + windows := map[tenant.ID]time.Duration{} + for id, store := range tenants.All() { + windows[id] = store.GapWindow() } - return window + return windows } // served reports whether the registry is serving tenant id: what the @@ -185,10 +174,11 @@ func perTenant[T any](tenants *settings.Registry, get func(*settings.Store) T) f // miss reads as DLQ on, not as the zero value perTenant would give: off lets // the worker drop a message it cannot read, and not knowing the tenant is no // reason to destroy its row. Parked, it survives until the tenant resolves. -// So a removed or rejected tenant's queued rows are parked under its own -// subject rather than left unacked for its return: an unacked row holds the -// ack floor, the sweeper stops purging, and the one shared stream fills -// toward mq.max_bytes_gb until every tenant's ingest answers 503. +// So a removed or rejected tenant's queued rows are parked in its own +// dead-letter queue rather than left unacked for its return: unacked, each +// would be redelivered every ack wait for as long as the tenant is away, and +// would hold the tenant's ack floor, so the sweeper could purge none of its +// queue past it. func dlqFor(tenants *settings.Registry) func(tenant.ID, string) bool { return func(id tenant.ID, table string) bool { store, ok := tenants.For(id) @@ -541,15 +531,23 @@ func (a *App) wireDedupe() error { } // wireMQ starts the MQ — the embedded NATS under data_dir/nats, the one -// place the implementation is chosen; everything after it sees mq.Broker. -// mq.max_bytes_gb is hot-reloadable: after each adoption the new budget is -// handed to the MQ, which owns how it is split across its queues and keeps -// them consistent (see mq.Broker.SetMaxBytes). +// place the implementation is chosen; everything after it sees mq.Broker — +// and hands it each served tenant's mq.max_bytes_gb, which opens that +// tenant's queue the first time. The budget is hot-reloadable: after every +// reload the registry applies, each served tenant's is handed over again, +// and the MQ owns how it is split across the tenant's queues and keeps them +// consistent (see mq.Broker.SetMaxBytes). A tenant no longer served keeps +// its queue at the budget it last had. A queue that cannot be opened or +// resized follows the registry's rule for the shape: a flat directory +// refuses boot, like every other store, and on a reload logs it, keeping the +// previous budget; a nested directory logs it at boot too, so it never costs +// the process — the tenant's ingest answers 503 until a reload opens its +// queue. The hook is registered before the boot apply, as the dedupe one is. func (a *App) wireMQ() error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker - broker, err := mq.NewEmbedded(dir, defaultSetting(a, (*settings.Store).MQMaxBytes)) + broker, err := mq.NewEmbedded(dir) if err != nil { config.LogStorageInitError("mq", dir, err) return fmt.Errorf("mq open: %w", err) @@ -570,17 +568,26 @@ func (a *App) wireMQ() error { // Rooted in the App's stop context, so a reload caught mid-hook by // SIGTERM gives up rather than holding the drain past // server.shutdown_timeout. - a.onDefaultAdopt(func() { - mb := defaultSetting(a, (*settings.Store).MQMaxBytes) - if mb == broker.MaxBytes() { - return - } - if err := broker.SetMaxBytes(a.stopCtx, mb); err != nil { - slog.Error("mq stream resize failed; the next reload retries", "error", err) - return + reconcile := func() error { + var errs []error + for id, store := range a.tenants.All() { + mb := store.MQMaxBytes() + if mb == broker.MaxBytes(id) { + continue + } + if err := broker.SetMaxBytes(a.stopCtx, id, mb); err != nil { + slog.Error("mq queue not reconciled with settings; the next reload retries", "tenant", id, "error", err) + errs = append(errs, fmt.Errorf("tenant %s: %w", id, err)) + continue + } + slog.Info("mq queue reconciled with settings", "tenant", id, "max_bytes_gb", mb>>30) } - slog.Info("mq stream limits reconciled with settings", "max_bytes_gb", mb>>30) - }) + return errors.Join(errs...) + } + a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile() }) + if err := reconcile(); err != nil && !a.tenants.Nested() { + return fmt.Errorf("mq open: %w", err) + } return nil } @@ -597,11 +604,11 @@ func (a *App) wireCache() error { } // wireSweeper adds the active sweeper — purges messages that are both -// written to ClickHouse and older than the SSE gap window (the longest -// stream.gap_window_minutes among the tenants served, re-read every sweep — -// see longestGapWindow). Runs every minute. +// written to ClickHouse and older than their tenant's SSE gap window (its own +// stream.gap_window_minutes, re-read every sweep — see gapWindows). Runs +// every minute. func (a *App) wireSweeper() { - sweeper := ingest.NewSweeper(a.mq, func() time.Duration { return longestGapWindow(a.tenants) }) + sweeper := ingest.NewSweeper(a.mq, func() map[tenant.ID]time.Duration { return gapWindows(a.tenants) }) a.add(component{name: "sweeper", run: func(ctx context.Context) error { sweeper.Start(ctx) return nil diff --git a/internal/ingest/sweeper.go b/internal/ingest/sweeper.go index 85383ead..367e22af 100644 --- a/internal/ingest/sweeper.go +++ b/internal/ingest/sweeper.go @@ -7,35 +7,34 @@ import ( "time" "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/tenant" ) // Sweeper implements the Active Sweeper pattern. It runs every minute and // asks the MQ to purge the ingest events that satisfy BOTH conditions: // - ACKed by the buffer consumer (written to ClickHouse) -// - Older than the gap window (no longer needed for SSE replay) +// - Older than their tenant's gap window (no longer needed for SSE replay) // -// This guarantees: healthy state keeps exactly gap_window of rolling data; -// ClickHouse down freezes purging; a catastrophic outage fills the queue to -// its byte budget and triggers backpressure (mq.ErrQueueFull). How the MQ -// finds the purge point is its own business (see mq.Purger). +// This guarantees: healthy state keeps exactly each tenant's gap_window of +// rolling data; ClickHouse down freezes purging; a catastrophic outage fills +// a tenant's queue to its byte budget and triggers backpressure +// (mq.ErrQueueFull). How the MQ finds the purge point is its own business +// (see mq.Purger). type Sweeper struct { purger mq.Purger - // gapWindow is the history to keep, read on every sweep so a reload of - // stream.gap_window_minutes applies from the next sweep without a - // restart. The ingest queue is one stream for every tenant and a purge - // is one bound over it, so in production this is the longest window - // among the tenants being served (internal/app's longestGapWindow); a - // tenant's own window follows once the streams are per tenant (#583 - // story 5b). - gapWindow func() time.Duration + // gapWindows is the history to keep for each tenant being served, read on + // every sweep so a reload of stream.gap_window_minutes applies from the + // next sweep without a restart. A tenant it does not name — one removed + // or rejected — keeps no history (mq.Purger.PurgeAcked). + gapWindows func() map[tenant.ID]time.Duration } -// NewSweeper creates the Active Sweeper. gapWindow is resolved per sweep. +// NewSweeper creates the Active Sweeper. gapWindows is resolved per sweep. // TODO: (future) need leader election or shared lock to only run one instance of the sweeper in clustered mode -func NewSweeper(purger mq.Purger, gapWindow func() time.Duration) *Sweeper { +func NewSweeper(purger mq.Purger, gapWindows func() map[tenant.ID]time.Duration) *Sweeper { return &Sweeper{ - purger: purger, - gapWindow: gapWindow, + purger: purger, + gapWindows: gapWindows, } } @@ -54,7 +53,13 @@ func (s *Sweeper) Start(ctx context.Context) { } func (s *Sweeper) sweep(ctx context.Context) { - _, err := s.purger.PurgeAcked(ctx, BufferConsumerName, time.Now().Add(-s.gapWindow())) + now := time.Now() + windows := s.gapWindows() + cutoffs := make(map[tenant.ID]time.Time, len(windows)) + for id, window := range windows { + cutoffs[id] = now.Add(-window) + } + _, err := s.purger.PurgeAcked(ctx, BufferConsumerName, cutoffs) if err != nil { if errors.Is(err, mq.ErrConsumerNotFound) { // Consumer may not exist yet if no messages have been ingested. diff --git a/internal/ingest/sweeper_test.go b/internal/ingest/sweeper_test.go index dfd4e02f..9b1eabab 100644 --- a/internal/ingest/sweeper_test.go +++ b/internal/ingest/sweeper_test.go @@ -7,6 +7,7 @@ import ( "time" "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -15,11 +16,11 @@ import ( // The purge-point arithmetic is the MQ's (internal/mq/purge_test.go); the // sweeper owns only when to ask and what window to ask for. -func TestSweep_AsksForTheBufferConsumerAndTheGapWindow(t *testing.T) { +func TestSweep_AsksForTheBufferConsumerAndEachTenantsGapWindow(t *testing.T) { t.Parallel() - gapWindow := 5 * time.Minute + windows := map[tenant.ID]time.Duration{"acme": 5 * time.Minute, "globex": time.Hour} purger := &testutil.MockPurger{Purged: true} - s := NewSweeper(purger, func() time.Duration { return gapWindow }) + s := NewSweeper(purger, func() map[tenant.ID]time.Duration { return windows }) before := time.Now() s.sweep(context.Background()) @@ -28,29 +29,35 @@ func TestSweep_AsksForTheBufferConsumerAndTheGapWindow(t *testing.T) { require.Len(t, purger.Calls, 1) call := purger.Calls[0] assert.Equal(t, BufferConsumerName, call.Consumer) - assert.False(t, call.OlderThan.Before(before.Add(-gapWindow)), "cutoff is now - gap window") - assert.False(t, call.OlderThan.After(after.Add(-gapWindow)), "cutoff is now - gap window") + require.Len(t, call.OlderThan, 2, "one cutoff per tenant served") + for id, window := range windows { + cutoff := call.OlderThan[id] + assert.False(t, cutoff.Before(before.Add(-window)), "%s: cutoff is now - its own gap window", id) + assert.False(t, cutoff.After(after.Add(-window)), "%s: cutoff is now - its own gap window", id) + } } -func TestSweep_RereadsTheGapWindowEverySweep(t *testing.T) { +func TestSweep_RereadsTheGapWindowsEverySweep(t *testing.T) { t.Parallel() - gapWindow := time.Minute + windows := map[tenant.ID]time.Duration{"acme": time.Minute} purger := &testutil.MockPurger{} - s := NewSweeper(purger, func() time.Duration { return gapWindow }) + s := NewSweeper(purger, func() map[tenant.ID]time.Duration { return windows }) s.sweep(context.Background()) - gapWindow = time.Hour // a settings reload + windows = map[tenant.ID]time.Duration{"acme": time.Hour, "globex": time.Minute} // a settings reload s.sweep(context.Background()) require.Len(t, purger.Calls, 2) - assert.Greater(t, purger.Calls[0].OlderThan.Sub(purger.Calls[1].OlderThan), 50*time.Minute) + assert.Greater(t, purger.Calls[0].OlderThan["acme"].Sub(purger.Calls[1].OlderThan["acme"]), 50*time.Minute) + assert.NotContains(t, purger.Calls[0].OlderThan, tenant.ID("globex")) + assert.Contains(t, purger.Calls[1].OlderThan, tenant.ID("globex"), "a tenant adopted since is named from the next sweep") } func TestSweep_ErrorsDoNotPanic(t *testing.T) { t.Parallel() for _, err := range []error{mq.ErrConsumerNotFound, errors.New("broker unavailable")} { purger := &testutil.MockPurger{Err: err} - s := NewSweeper(purger, func() time.Duration { return time.Minute }) + s := NewSweeper(purger, func() map[tenant.ID]time.Duration { return map[tenant.ID]time.Duration{"acme": time.Minute} }) s.sweep(context.Background()) assert.Len(t, purger.Calls, 1) } @@ -62,7 +69,7 @@ func TestSweep_ErrorsDoNotPanic(t *testing.T) { func TestStart_ContextCancellation(t *testing.T) { t.Parallel() - s := NewSweeper(&testutil.MockPurger{}, func() time.Duration { return 5 * time.Minute }) + s := NewSweeper(&testutil.MockPurger{}, func() map[tenant.ID]time.Duration { return nil }) ctx, cancel := context.WithCancel(context.Background()) cancel() // Cancel immediately. diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index 618b2b35..9af5b1ef 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -116,10 +116,12 @@ const ( // maxAckPending, and ackWait > defaultMaxWait + CH flush (else in-flight // messages are redelivered mid-processing → duplicate inserts). const ( - // Server-side cap on unacked messages; suspends delivery when hit (backpressure). + // Server-side cap on a tenant's unacked messages; suspends that tenant's + // delivery when hit (backpressure), and no other tenant's. maxAckPending = 10_000 // TODO: raise if NATS delivery becomes the bottleneck - // Client prefetch buffer in front of msgChan (was the implicit jetstream default). + // Client prefetch buffer in front of msgChan (was the implicit jetstream + // default), shared by the tenants' queues (mq.Consumer.Consume). pullMaxMessages = 500 // Redelivery timeout. 60s ≈ 5s batch + ~30s HTTP timeout + margin. @@ -219,8 +221,9 @@ func waitOrDeadline(ctx context.Context, wg *sync.WaitGroup) error { } } -// dispatchLoop owns the single JetStream consumer and fans every message out to -// a tableLoop per tenant table (lazily spawned on first sight of one). It does +// dispatchLoop owns the one consumer — held on every tenant's queue — and fans +// every message out to a tableLoop per tenant table (lazily spawned on first +// sight of one). It does // no batching itself — it parses just enough to route — so a low-volume table // can never strand another table's rows behind a shared timer. It is the ONLY // goroutine that watches ctx; tableLoops stop via channel-close, which gives a @@ -230,11 +233,13 @@ func (w *IngestWorker) dispatchLoop(ctx context.Context, cons mq.Consumer) { msgChan := make(chan *mq.Message, w.maxBatch*2) - // Pull consumer with a push-like callback (the client prefetches pullMaxMessages). - // Hand off to msgChan only, so the consume goroutine never blocks on flush work. + // Pull consumer with a push-like callback (the client prefetches pullMaxMessages, + // shared by the tenants' queues). It runs on one delivery goroutine per tenant, + // so the handoff is a channel send, safe from all of them at once. Hand off to + // msgChan only, so a consume goroutine never blocks on flush work. // The handoff also watches ctx: stop (deferred below) does not wait for a // delivery already in the handler, so once this loop has stopped draining - // msgChan a full channel would otherwise pin the client's delivery goroutine + // msgChan a full channel would otherwise pin a delivery goroutine // forever. A message dropped here is unacked and simply redelivered. stop, deliveryEnded, err := cons.Consume(func(msg *mq.Message) { select { diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index c3a0988e..7fe9130e 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -122,9 +122,7 @@ func TestStartIngestWorker_Validation(t *testing.T) { { name: "nil cache", setup: func(t *testing.T) (Queue, cache.Cache) { - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) return emb, nil }, wantErrSub: "cache is nil", @@ -154,9 +152,7 @@ func TestStartIngestWorker_EndToEnd(t *testing.T) { t.Parallel() // ── Embedded MQ ── - emb, err := mq.NewEmbedded(t.TempDir(), 4*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 4*1024*1024) // ── ClickHouse stub: capture each request body, return 200 ── var ( @@ -247,9 +243,7 @@ func TestStartIngestWorker_EndToEnd(t *testing.T) { func TestStartIngestWorker_StopFunc_RespectsShutdownDeadline(t *testing.T) { t.Parallel() - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) // ClickHouse stub that blocks until we say go — keeps the worker's // flush goroutine alive past the stop call. @@ -298,9 +292,7 @@ func TestStartIngestWorker_StopFunc_RespectsShutdownDeadline(t *testing.T) { func TestStartIngestWorker_StopFunc_CleanShutdown(t *testing.T) { t.Parallel() - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) // chURL is never dialed: with no messages there is no flush, so a dummy // host/port is fine. @@ -1094,9 +1086,7 @@ func TestDispatchLoop_PerTableBatching_NoCrossTableContamination(t *testing.T) { batchB = maxBatch ) - emb, err := mq.NewEmbedded(t.TempDir(), 8*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 8*1024*1024) // CH stub: count rows (newlines in the JSONCompactEachRow body) per target table. var ( @@ -1187,9 +1177,7 @@ func TestDispatchLoop_PartialBatchWaitsForOwnTrigger(t *testing.T) { total = 4 // 3 → one full batch on the size trigger; 1 leftover ) - emb, err := mq.NewEmbedded(t.TempDir(), 8*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 8*1024*1024) // CH stub counts rows and sleeps briefly, so the 4th row is reliably buffered // before the first (3-row) flush completes — that's when the old code would @@ -1791,9 +1779,7 @@ func TestDispatchLoop_BatchesPerTenantTable(t *testing.T) { t.Parallel() const maxBatch = 2 - emb, err := mq.NewEmbedded(t.TempDir(), 8*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 8*1024*1024, "acme", "globex") // CH stub: record each INSERT's body under the database it named. var ( diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 4219c34d..dbe35fa5 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -5,12 +5,16 @@ import ( "errors" "fmt" "log/slog" + "maps" + "math" + "slices" "strings" "sync" "sync/atomic" "time" "github.com/Wave-RF/WaveHouse/internal/observability" + "github.com/Wave-RF/WaveHouse/internal/tenant" natsserver "github.com/nats-io/nats-server/v2/server" "github.com/nats-io/nats.go" "github.com/nats-io/nats.go/jetstream" @@ -44,45 +48,73 @@ func (slogNATSLogger) Tracef(format string, v ...any) { slog.Debug(fmt.Sprintf(format, v...), "component", "nats") } -// EmbeddedNATS runs an in-process NATS server with JetStream. +// EmbeddedNATS runs an in-process NATS server with JetStream, and gives each +// tenant a queue of its own: an ingest stream and a dead-letter stream +// (subject.go names them), each with its own byte cap, and the durable +// consumers on the ingest one. Nothing outside this package sees that layout. type EmbeddedNATS struct { server *natsserver.Server conn *nats.Conn js jetstream.JetStream - limitMu sync.Mutex - maxBytes int64 // the ingest stream cap both streams were last reconciled to + // mu guards queues and consumers, and serializes opening or resizing a + // tenant's queue with registering a consumer, so a queue opened while a + // consumer registers is never missed by it. It is held across the + // JetStream calls that open or resize a queue. + mu sync.Mutex + queues map[tenant.ID]*tenantQueue + // consumers are the durable consumers held on every tenant's queue, each + // joined to a queue as it opens. + consumers []*fanIn +} + +// tenantQueue is what the broker knows of one tenant's queue. +type tenantQueue struct { + // ingest and dlq report whether each of the tenant's streams exists. + ingest, dlq bool + // maxBytes is the budget last applied in full (MaxBytes); asked is the + // budget last asked for, which a publish or park that finds a stream + // missing opens it at. Both are read back from the ingest stream at boot, + // so a tenant no longer served keeps the budget it last had. + maxBytes, asked int64 } // EmbeddedNATS is the one implementation of every mq interface. var _ Broker = (*EmbeddedNATS)(nil) const ( - // dlqShare is the DLQ stream's slice of the byte budget: a tenth of the - // ingest stream's cap. + // dlqShare is a tenant's dead-letter stream's slice of its byte budget: a + // tenth of the ingest stream's cap. dlqShare = 10 - // resizeTimeout bounds the JetStream calls SetMaxBytes makes to apply a - // new cap — both streams share it. A settings reload holds the store's - // lock while its hooks run, so an in-process JetStream call that never - // returns would otherwise block every later reload. + // resizeTimeout bounds the JetStream calls SetMaxBytes makes to open a + // tenant's queue or apply a new cap to it — both streams share it. A + // settings reload holds the store's lock while its hooks run, so an + // in-process JetStream call that never returns would otherwise block every + // later reload. resizeTimeout = 10 * time.Second - // rollbackTimeout is the undo's own budget when the DLQ resize fails: - // in-process JetStream fails by stalling rather than erroring, so the - // likely cause is that resizeTimeout has just run out, and an undo on + // rollbackTimeout is the undo's own budget when the dead-letter resize + // fails: in-process JetStream fails by stalling rather than erroring, so + // the likely cause is that resizeTimeout has just run out, and an undo on // that context would fail without touching the stream. SetMaxBytes runs // for at most the sum of the two. rollbackTimeout = 5 * time.Second ) -// NewEmbedded starts an embedded NATS server with JetStream enabled and -// both streams in place: the ingest stream capped at maxBytes and the DLQ -// stream at a tenth of it. The DLQ stream is always present — an empty -// limits-policy stream costs nothing, and whether a poison row lands on it is -// the ingest worker's decision at the moment of the failure. -// The server logs through slog's default logger. The stream names are fixed -// (see subject.go) — the embedded server is private to this process, so -// there's no namespacing to do. -func NewEmbedded(storeDir string, maxBytes int64) (*EmbeddedNATS, error) { +// errNoQueue is why a publish or park finds no queue it can open: no budget +// has been asked for the tenant yet (see SetMaxBytes). Publish reports it as +// ErrQueueFull. +var errNoQueue = errors.New("no queue is open for it yet") + +// NewEmbedded starts an embedded NATS server with JetStream over storeDir and +// takes stock of the tenants' queues already there: a consumer created later +// is held on every one of them, those of tenants no longer served included, +// whose queued rows still have to reach the ingest worker. The pair of streams +// an earlier build kept for every tenant together is deleted, since its +// subjects overlap every tenant's; the events it held are not carried over. A +// tenant's queue is opened by SetMaxBytes, the first time its budget is +// applied, or by a publish or park that finds it missing, at the budget last +// asked for it. The server logs through slog's default logger. +func NewEmbedded(storeDir string) (*EmbeddedNATS, error) { opts := &natsserver.Options{ DontListen: true, JetStream: true, @@ -93,6 +125,18 @@ func NewEmbedded(storeDir string, maxBytes int64) (*EmbeddedNATS, error) { // channel" panic) and os.Exit(0)s past its cleanup. WaveHouse owns // the lifecycle; Close() shuts the server down. See #287. NoSigs: true, + // JetStream counts every stream's byte cap as reserved disk and + // refuses a stream once the caps together pass this limit — by + // default 75% of the free disk at boot. A tenant's mq.max_bytes_gb + // caps that tenant's queue and nothing else; what the tenants' caps + // add up to against the disk is #138's to decide, not a limit the + // server enforces on the side, so its own is set out of reach. Half + // the int64 range, not all of it: the server subtracts its count of + // reserved bytes from this limit, and a stream whose store fails to + // open releases a reservation it never made (nats-server 2.14.6), so + // the count can fall below zero — at the top of the range that + // subtraction overflows, and every stream after it is refused. + JetStreamMaxStore: math.MaxInt64 / 2, } ns, err := natsserver.NewServer(opts) @@ -119,114 +163,285 @@ func NewEmbedded(storeDir string, maxBytes int64) (*EmbeddedNATS, error) { return nil, fmt.Errorf("jetstream new: %w", err) } - if _, err := js.CreateOrUpdateStream(context.Background(), ingestStreamConfig(maxBytes)); err != nil { - nc.Close() - ns.Shutdown() - return nil, fmt.Errorf("create stream: %w", err) + e := &EmbeddedNATS{server: ns, conn: nc, js: js, queues: map[tenant.ID]*tenantQueue{}} + if err := e.takeStock(context.Background()); err != nil { + _ = e.Close() + return nil, err } - if _, err := js.CreateOrUpdateStream(context.Background(), dlqStreamConfig(maxBytes/dlqShare)); err != nil { - nc.Close() - ns.Shutdown() - return nil, fmt.Errorf("create dlq stream: %w", err) + return e, nil +} + +// takeStock deletes the pair of streams an earlier build kept for every +// tenant together, then records every tenant stream on disk, with the budget +// its ingest stream last had. +func (e *EmbeddedNATS) takeStock(ctx context.Context) error { + for _, name := range []string{legacyIngestStream, legacyDLQStream} { + if err := e.deleteLegacy(ctx, name); err != nil { + return err + } + } + streams := e.js.ListStreams(ctx) + for info := range streams.Info() { + name := info.Config.Name + if id, ok := streamTenant(ingestStreamPrefix, name); ok { + q := e.queue(id) + q.ingest = true + q.maxBytes, q.asked = info.Config.MaxBytes, info.Config.MaxBytes + } else if id, ok := streamTenant(dlqStreamPrefix, name); ok { + e.queue(id).dlq = true + } } + if err := streams.Err(); err != nil { + return fmt.Errorf("list streams: %w", err) + } + return nil +} - return &EmbeddedNATS{server: ns, conn: nc, js: js, maxBytes: maxBytes}, nil +// deleteLegacy deletes one stream of the pair an earlier build kept for every +// tenant together, logging what it held; one that is not there is nothing to +// do. +func (e *EmbeddedNATS) deleteLegacy(ctx context.Context, name string) error { + s, err := e.js.Stream(ctx, name) + if errors.Is(err, jetstream.ErrStreamNotFound) { + return nil + } + if err != nil { + return fmt.Errorf("look up stream %s: %w", name, err) + } + held := s.CachedInfo().State.Msgs + if err := e.js.DeleteStream(ctx, name); err != nil { + return fmt.Errorf("delete stream %s: %w", name, err) + } + slog.Warn("mq: deleted the stream an earlier build kept for every tenant together; its messages are not carried over", + "component", "nats", "stream", name, "messages", held) + return nil +} + +// queue returns what the broker knows of tenant id's queue, recording the +// tenant first if it knows nothing. Under e.mu (or before e is shared). +func (e *EmbeddedNATS) queue(id tenant.ID) *tenantQueue { + q := e.queues[id] + if q == nil { + q = &tenantQueue{} + e.queues[id] = q + } + return q } -// ingestStreamConfig is the WAVEHOUSE stream. LimitsPolicy: standard +// ingestTenants lists the tenants whose ingest stream exists, in id order. +// Under e.mu. +func (e *EmbeddedNATS) ingestTenants() []tenant.ID { + var ids []tenant.ID + for id, q := range e.queues { + if q.ingest { + ids = append(ids, id) + } + } + slices.Sort(ids) + return ids +} + +// ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard // append-only log; the Active Sweeper handles message purging. MaxBytes caps -// disk usage to protect the shared ClickHouse/NATS disk. DiscardNew rejects -// new messages when full, propagating backpressure to the upstream API. -func ingestStreamConfig(maxBytes int64) jetstream.StreamConfig { +// the tenant's share of the disk. DiscardNew rejects new messages when full, +// propagating backpressure to the upstream API — for this tenant alone. +func ingestStreamConfig(id tenant.ID, maxBytes int64) jetstream.StreamConfig { return jetstream.StreamConfig{ - Name: ingestStream, - Subjects: []string{ingestAll}, + Name: ingestStreamName(id), + Subjects: []string{tenantSubjects(ingestPrefix, id)}, Retention: jetstream.LimitsPolicy, MaxBytes: maxBytes, Discard: jetstream.DiscardNew, } } -// dlqStreamConfig is the WAVEHOUSE_DLQ stream. DiscardOld: a full DLQ drops -// its oldest parked rows rather than refusing new ones — backpressure belongs -// to the ingest stream, not the dead-letter one. -func dlqStreamConfig(maxBytes int64) jetstream.StreamConfig { +// dlqStreamConfig is tenant id's dead-letter stream. DiscardOld: a full one +// drops its oldest parked rows rather than refusing new ones — backpressure +// belongs to the ingest stream, not the dead-letter one. +func dlqStreamConfig(id tenant.ID, maxBytes int64) jetstream.StreamConfig { return jetstream.StreamConfig{ - Name: dlqStream, - Subjects: []string{dlqAll}, + Name: dlqStreamName(id), + Subjects: []string{tenantSubjects(dlqPrefix, id)}, Retention: jetstream.LimitsPolicy, MaxBytes: maxBytes, Discard: jetstream.DiscardOld, } } -// MaxBytes reports the ingest stream cap both streams were last reconciled to -// (by NewEmbedded, then by each successful SetMaxBytes). -func (e *EmbeddedNATS) MaxBytes() int64 { - e.limitMu.Lock() - defer e.limitMu.Unlock() - return e.maxBytes +// MaxBytes reports the budget tenant id's queue was last given in full (by +// SetMaxBytes, or read back from disk at boot), 0 when it has none. +func (e *EmbeddedNATS) MaxBytes(id tenant.ID) int64 { + e.mu.Lock() + defer e.mu.Unlock() + if q := e.queues[id]; q != nil { + return q.maxBytes + } + return 0 } -// SetMaxBytes applies a new byte budget to both streams in place (the -// hot-reloadable mq.max_bytes_gb): the ingest stream takes maxBytes and the -// DLQ stream a tenth of it. JetStream applies a limit change to a live stream -// without touching its messages: growing takes effect immediately; shrinking -// the ingest stream below its current size makes DiscardNew refuse new -// publishes until the worker drains it — nothing buffered is dropped. +// SetMaxBytes applies tenant id's byte budget (its hot-reloadable +// mq.max_bytes_gb) to its queue: the ingest stream takes maxBytes and the +// dead-letter stream a tenth of it. A tenant with no queue yet has one opened, +// its dead-letter stream first, so no row is queued that could not be parked, +// and every registered consumer joins it. No other tenant's queue is touched. +// +// JetStream applies a limit change to a live stream without touching its +// messages: growing takes effect immediately; shrinking the ingest stream +// below its current size makes DiscardNew refuse new publishes until the +// worker drains it — nothing buffered is dropped. The dead-letter stream is +// DiscardOld, which would delete its oldest parked rows to fit a smaller cap, +// so it is never capped below the bytes it holds (#532): it keeps what it +// has, and that is logged. // -// The pair moves together where it can. If the DLQ update fails after the -// ingest one succeeded, the ingest resize is undone so the 10:1 pair stays at +// The pair moves together where it can. If the dead-letter update fails after +// the ingest one succeeded, the ingest resize is undone so the pair stays at // the previous budget, and the next call retries both. Safe in that direction // — the ingest stream is DiscardNew, so shrinking it back drops nothing // stored. The undo is best effort: if it fails too, the ingest stream stays at -// the new limit and the DLQ at the previous, and the error says so. On any -// error MaxBytes keeps reporting the previous budget, so a later call with the -// new budget reapplies both. +// the new limit and the dead-letter one at the previous, and the error says +// so. On any error MaxBytes keeps reporting the previous budget, so a later +// call with the new budget reapplies both. // // The JetStream calls are bounded by resizeTimeout, plus rollbackTimeout for // the undo, both rooted in ctx. That is deliberate: ctx is the process's stop // context, so a reload caught mid-hook by a stop gives up — undo included — // rather than holding the drain past server.shutdown_timeout. A cancellation // between the two updates is therefore the one way to leave the pair split, -// and only for the rest of a process that is exiting: the next boot -// reconciles both streams from the adopted settings. -func (e *EmbeddedNATS) SetMaxBytes(ctx context.Context, maxBytes int64) error { - e.limitMu.Lock() - defer e.limitMu.Unlock() - if maxBytes == e.maxBytes { +// and only for the rest of a process that is exiting: the next boot applies +// the adopted settings to it again. +func (e *EmbeddedNATS) SetMaxBytes(ctx context.Context, id tenant.ID, maxBytes int64) error { + if _, err := tenant.Parse(string(id)); err != nil { + return fmt.Errorf("tenant: %w", err) + } + e.mu.Lock() + defer e.mu.Unlock() + q := e.queue(id) + q.asked = maxBytes + if q.ingest && q.dlq && maxBytes == q.maxBytes { return nil } + return e.apply(ctx, id, q, maxBytes) +} +// apply brings tenant id's queue to maxBytes: opening it when its ingest +// stream is missing, resizing it otherwise (see SetMaxBytes). Under e.mu. +func (e *EmbeddedNATS) apply(ctx context.Context, id tenant.ID, q *tenantQueue, maxBytes int64) error { resizeCtx, cancel := context.WithTimeout(ctx, resizeTimeout) defer cancel() - if _, err := e.js.UpdateStream(resizeCtx, ingestStreamConfig(maxBytes)); err != nil { + if !q.ingest { + if err := e.applyDLQ(resizeCtx, id, q, maxBytes); err != nil { + return err + } + if _, err := e.js.CreateOrUpdateStream(resizeCtx, ingestStreamConfig(id, maxBytes)); err != nil { + return fmt.Errorf("open ingest stream: %w", err) + } + q.ingest, q.maxBytes = true, maxBytes + for _, f := range e.consumers { + if err := f.join(resizeCtx, id); err != nil { + f.fail(fmt.Errorf("tenant %s: %w: join its queue: %w", id, ErrDeliveryEnded, err)) + } + } + return nil + } + if _, err := e.js.UpdateStream(resizeCtx, ingestStreamConfig(id, maxBytes)); err != nil { return fmt.Errorf("resize ingest stream: %w", err) } - if _, err := e.js.CreateOrUpdateStream(resizeCtx, dlqStreamConfig(maxBytes/dlqShare)); err != nil { - // The undo runs on its own budget, not the one the DLQ call has - // likely just exhausted. + if err := e.applyDLQ(resizeCtx, id, q, maxBytes); err != nil { + // The undo runs on its own budget, not the one the dead-letter call + // has likely just exhausted. rollbackCtx, cancelRollback := context.WithTimeout(ctx, rollbackTimeout) defer cancelRollback() - if _, rollbackErr := e.js.UpdateStream(rollbackCtx, ingestStreamConfig(e.maxBytes)); rollbackErr != nil { - return fmt.Errorf("resize dlq stream: %w (ingest stream rollback failed, so it stays at the new limit and the dlq at the previous: %w)", err, rollbackErr) + if _, rollbackErr := e.js.UpdateStream(rollbackCtx, ingestStreamConfig(id, q.maxBytes)); rollbackErr != nil { + return fmt.Errorf("%w (ingest stream rollback failed, so it stays at the new limit and the dlq at the previous: %w)", err, rollbackErr) } - return fmt.Errorf("resize dlq stream: %w (ingest stream restored to the previous limit)", err) + return fmt.Errorf("%w (ingest stream restored to the previous limit)", err) } - e.maxBytes = maxBytes + q.maxBytes = maxBytes return nil } -// Publish stores data on topic's ingest subject. A topic without a valid -// tenant is refused before anything is sent (see subject). A stream at its -// byte budget (DiscardNew) refuses the publish; that is reported as -// ErrQueueFull. +// applyDLQ gives tenant id's dead-letter stream a tenth of maxBytes, creating +// it when it is missing, but never caps it below the bytes it holds: those +// stay, the cap is what they take, and the stream then drops its oldest row +// to make room for each new one, as any full dead-letter stream does. Under +// e.mu. +func (e *EmbeddedNATS) applyDLQ(ctx context.Context, id tenant.ID, q *tenantQueue, maxBytes int64) error { + limit := maxBytes / dlqShare + verb := "resize" + s, err := e.js.Stream(ctx, dlqStreamName(id)) + switch { + case errors.Is(err, jetstream.ErrStreamNotFound): + verb = "open" + case err != nil: + return fmt.Errorf("dlq stream info: %w", err) + default: + // A stream's size fits an int64 as its cap does; the bound is + // checked rather than assumed. + if held := s.CachedInfo().State.Bytes; held <= math.MaxInt64 && int64(held) > limit { + slog.Warn("mq: dead-letter queue kept at what it holds rather than shrunk to its budget, so no parked row is deleted", + "component", "nats", "tenant", id, "held_bytes", held, "budget_bytes", limit) + limit = int64(held) + } + } + if _, err := e.js.CreateOrUpdateStream(ctx, dlqStreamConfig(id, limit)); err != nil { + return fmt.Errorf("%s dlq stream: %w", verb, err) + } + q.dlq = true + return nil +} + +// reopen opens tenant id's queue at the budget last asked for it, for a +// publish or park that found one of its streams missing. errNoQueue when no +// budget has been asked for the tenant yet: a reload can make a tenant +// resolvable an instant before its budget arrives. +func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { + e.mu.Lock() + defer e.mu.Unlock() + q := e.queues[id] + if q == nil || q.asked == 0 { + return fmt.Errorf("tenant %s: %w", id, errNoQueue) + } + // What is missing is asked of JetStream rather than read off the flags, + // which may still say the stream the publish just missed exists — or it + // may be back already, opened by a caller that held mu first. + for _, name := range []string{ingestStreamName(id), dlqStreamName(id)} { + _, err := e.js.Stream(ctx, name) + switch { + case errors.Is(err, jetstream.ErrStreamNotFound): + if name == ingestStreamName(id) { + q.ingest = false + } else { + q.dlq = false + } + case err != nil: + return fmt.Errorf("stream info: %w", err) + } + } + if q.ingest && q.dlq { + return nil + } + return e.apply(ctx, id, q, q.asked) +} + +// Publish stores data on topic's ingest subject, in its tenant's queue. A +// topic without a valid tenant is refused before anything is sent (see +// subject). A tenant with no queue has one opened at the budget last asked +// for it (see SetMaxBytes). A queue that cannot be opened — none asked for +// yet, or JetStream refused it — and a queue at its byte budget (DiscardNew) +// are reported as ErrQueueFull: either way the tenant's queue takes nothing +// now, and a retry is the caller's answer. func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, opts ...PublishOpt) error { subj, err := subject(ingestPrefix, topic) if err != nil { return err } err = e.publish(ctx, subj, data, opts) + if errors.Is(err, jetstream.ErrNoStreamResponse) { + if openErr := e.reopen(ctx, topic.Tenant); openErr != nil { + return fmt.Errorf("%w: %w", ErrQueueFull, openErr) + } + err = e.publish(ctx, subj, data, opts) + } if err != nil && strings.Contains(err.Error(), "maximum bytes exceeded") { // The server reports a full store as a generic store failure whose // text is the only thing that names the cause. @@ -235,12 +450,23 @@ func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, op return err } -// DeadLetter stores msg's data on its topic's DLQ subject — the subject it -// arrived on with the ingest prefix swapped for the DLQ one, nothing decoded -// or re-encoded. The DLQ stream is DiscardOld, so a full DLQ drops its oldest -// parked rows rather than refusing. +// DeadLetter stores msg's data on its topic's dead-letter subject, in its +// tenant's queue — the subject it arrived on with the ingest prefix swapped +// for the dead-letter one, nothing decoded or re-encoded. The dead-letter +// stream is DiscardOld, so a full one drops its oldest parked rows rather than +// refusing. A dead-letter stream found missing is opened again with its +// tenant's queue, as Publish does. func (e *EmbeddedNATS) DeadLetter(ctx context.Context, msg *Message, opts ...PublishOpt) error { - return e.publish(ctx, dlqPrefix+msg.topicKey, msg.Data, opts) + subj := dlqPrefix + msg.topicKey + err := e.publish(ctx, subj, msg.Data, opts) + if errors.Is(err, jetstream.ErrNoStreamResponse) { + if id, ok := keyTenant(msg.topicKey); ok { + if err = e.reopen(ctx, id); err == nil { + err = e.publish(ctx, subj, msg.Data, opts) + } + } + } + return err } func (e *EmbeddedNATS) publish(ctx context.Context, subj string, data []byte, opts []PublishOpt) error { @@ -255,7 +481,10 @@ func (e *EmbeddedNATS) publish(ctx context.Context, subj string, data []byte, op observability.InjectHeaders(ctx, headers) msg.Header = nats.Header(headers) - _, err := e.js.PublishMsg(ctx, msg) + // No retry on "no responders": in-process, that only ever means no + // stream holds the subject — a tenant with no queue, which the callers + // open rather than wait out. + _, err := e.js.PublishMsg(ctx, msg, jetstream.WithRetryAttempts(0)) return err } @@ -275,58 +504,170 @@ func wrapMsg(ctx context.Context, m jetstream.Msg) *Message { ) } +// Subscribe holds a durable explicit-ack consumer named consumerName on every +// tenant's queue, those opened later included, and delivers each message to +// handler with the trace context its headers carry, until ctx is done. A +// tenant's queue that cannot be joined when it opens is logged: its events +// reach handler from the next boot. func (e *EmbeddedNATS) Subscribe(ctx context.Context, consumerName string, handler func(msg *Message) error) error { - cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStream, jetstream.ConsumerConfig{ - Durable: consumerName, - FilterSubject: ingestAll, - AckPolicy: jetstream.AckExplicitPolicy, - }) - if err != nil { + f := e.newFanIn(ctx, jetstream.ConsumerConfig{Durable: consumerName, AckPolicy: jetstream.AckExplicitPolicy}) + f.fail = func(err error) { + slog.Error("mq: a tenant's events do not reach this consumer until the next boot", "component", "nats", "consumer", consumerName, "error", err) + } + if err := e.register(ctx, f); err != nil { return fmt.Errorf("create consumer: %w", err) } - - cctx, err := cons.Consume(func(m jetstream.Msg) { + stop, err := f.start(func(m jetstream.Msg) { msg := wrapMsg(observability.ExtractHeaders(ctx, m.Headers()), m) if err := handler(msg); err != nil { _ = msg.Nak() } - }) + }, 0, false) if err != nil { return fmt.Errorf("consume: %w", err) } go func() { <-ctx.Done() - cctx.Stop() + stop() }() return nil } // CreateConsumer creates or updates a durable explicit-ack pull consumer on -// the ingest stream. ctx becomes every delivered Message.Ctx (see -// ConsumerManager); it does not stop delivery — Consumer.Consume's stop does. +// every tenant's queue, and joins each queue opened later. ctx becomes every +// delivered Message.Ctx (see ConsumerManager); it does not stop delivery — +// Consumer.Consume's stop does. func (e *EmbeddedNATS) CreateConsumer(ctx context.Context, cfg ConsumerConfig) (Consumer, error) { - cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStream, jetstream.ConsumerConfig{ - Durable: cfg.Durable, - FilterSubject: ingestAll, - AckPolicy: jetstream.AckExplicitPolicy, - AckWait: cfg.AckWait, - MaxAckPending: cfg.MaxAckPending, - }) - if err != nil { + c := &workerConsumer{ + fanIn: e.newFanIn(ctx, jetstream.ConsumerConfig{ + Durable: cfg.Durable, + AckPolicy: jetstream.AckExplicitPolicy, + AckWait: cfg.AckWait, + MaxAckPending: cfg.MaxAckPending, + }), + failed: make(chan error, 1), + } + c.fail = func(err error) { + // Exactly one error, and nothing once stop has been called. + if c.stopped.Load() { + return + } + select { + case c.failed <- err: + default: + } + } + if err := e.register(ctx, c.fanIn); err != nil { return nil, fmt.Errorf("create consumer: %w", err) } - return &jsConsumer{cons: cons, ctx: ctx}, nil + return c, nil +} + +// newFanIn is a fanIn over cfg, not yet holding any durable; the caller sets +// its fail and registers it. +func (e *EmbeddedNATS) newFanIn(ctx context.Context, cfg jetstream.ConsumerConfig) *fanIn { + return &fanIn{ + e: e, + ctx: ctx, + cfg: cfg, + handles: map[tenant.ID]jetstream.Consumer{}, + running: map[tenant.ID]jetstream.ConsumeContext{}, + } } -// jsConsumer is the Consumer over a JetStream pull consumer. -type jsConsumer struct { - cons jetstream.Consumer - ctx context.Context // each delivered Message.Ctx (see ConsumerManager) +// register holds f's durable on every tenant's queue there is and registers +// f, so every queue opened from here on is joined too. +func (e *EmbeddedNATS) register(ctx context.Context, f *fanIn) error { + e.mu.Lock() + defer e.mu.Unlock() + for _, id := range e.ingestTenants() { + if err := f.join(ctx, id); err != nil { + return fmt.Errorf("tenant %s: %w", id, err) + } + } + e.consumers = append(e.consumers, f) + return nil +} + +// unregister stops joining f to the queues that open from here on. Under +// e.mu. +func (e *EmbeddedNATS) unregister(f *fanIn) { + e.consumers = slices.DeleteFunc(e.consumers, func(c *fanIn) bool { return c == f }) } -func (c *jsConsumer) Consume(handler func(msg *Message), prefetch int) (func(), <-chan error, error) { +// fanIn is one durable consumer held on every tenant's ingest stream — the +// ingest worker's (CreateConsumer) or the hub bridge's (Subscribe) — +// delivering them all into one handler: each tenant's messages on a +// goroutine of their own, so a tenant's arrive in order and different +// tenants' concurrently, and a handler blocked on one tenant holds back that +// tenant alone. Its fields are guarded by e.mu, bar stopped. +type fanIn struct { + e *EmbeddedNATS + ctx context.Context // each delivered Message.Ctx (CreateConsumer), or where Subscribe extracts trace context into + cfg jetstream.ConsumerConfig + + // fail reports a tenant's delivery that ended on its own, or a queue that + // could not be joined when it opened. + fail func(error) + + // handles is the durable on each tenant's ingest stream; running, the + // delivery started on each once deliver is set. + handles map[tenant.ID]jetstream.Consumer + running map[tenant.ID]jetstream.ConsumeContext + deliver func(jetstream.Msg) + // prefetch is the fetch-ahead asked for across the tenants together; 0 + // leaves each tenant the client default. + prefetch int + // watch reports a delivery that ends on its own through fail. + watch bool + stopped atomic.Bool +} + +// join holds f's durable on tenant id's ingest stream — looked up first, and +// created or updated only when missing or configured otherwise, so a boot +// over thousands of queues writes nothing it need not — and starts delivery +// on it when f is delivering. Under e.mu. +func (f *fanIn) join(ctx context.Context, id tenant.ID) error { + stream := ingestStreamName(id) + c, err := f.e.js.Consumer(ctx, stream, f.cfg.Durable) + if err != nil || !sameConsumer(c.CachedInfo().Config, f.cfg) { + if c, err = f.e.js.CreateOrUpdateConsumer(ctx, stream, f.cfg); err != nil { + return err + } + } + f.handles[id] = c + if f.deliver == nil || f.stopped.Load() { + return nil + } + return f.run(id) +} + +// sameConsumer reports whether a durable holds the fields this package sets; +// a zero field in want is the server's default, whatever that resolved to. +func sameConsumer(have, want jetstream.ConsumerConfig) bool { + return have.AckPolicy == want.AckPolicy && + have.FilterSubject == want.FilterSubject && + (want.AckWait == 0 || have.AckWait == want.AckWait) && + (want.MaxAckPending == 0 || have.MaxAckPending == want.MaxAckPending) +} + +// share is one tenant's part of the fetch-ahead: the total spread over the +// tenants' queues, at least one each. 0 leaves the client default. Under +// e.mu. +func (f *fanIn) share() int { + if f.prefetch <= 0 { + return 0 + } + return max(1, f.prefetch/max(1, len(f.handles))) +} + +// run starts delivery from tenant id's durable, once. Under e.mu. +func (f *fanIn) run(id tenant.ID) error { + if _, ok := f.running[id]; ok { + return nil + } // The client reports what goes wrong after Consume returns only through // this handler, never through Consume's own error. It calls it for // passing conditions too (a missed heartbeat, a leadership change) and @@ -339,37 +680,80 @@ func (c *jsConsumer) Consume(handler func(msg *Message), prefetch int) (func(), opts := []jetstream.PullConsumeOpt{ jetstream.ConsumeErrHandler(func(_ jetstream.ConsumeContext, err error) { lastErr.Store(&err) - slog.Warn("mq: consumer reported an error", "component", "nats", "error", err) + slog.Warn("mq: consumer reported an error", "component", "nats", "tenant", id, "error", err) }), } - if prefetch > 0 { - opts = append(opts, jetstream.PullMaxMessages(prefetch)) + if n := f.share(); n > 0 { + opts = append(opts, jetstream.PullMaxMessages(n)) } - cctx, err := c.cons.Consume(func(m jetstream.Msg) { - handler(wrapMsg(c.ctx, m)) - }, opts...) + cctx, err := f.handles[id].Consume(f.deliver, opts...) if err != nil { - return nil, nil, fmt.Errorf("consume: %w", err) + return err + } + f.running[id] = cctx + if !f.watch { + return nil } - - var stopped atomic.Bool - failed := make(chan error, 1) go func() { <-cctx.Closed() - if stopped.Load() { + if f.stopped.Load() { return } - if reason := lastErr.Load(); reason != nil { - failed <- fmt.Errorf("%w: %w", ErrDeliveryEnded, *reason) - return + reason := ErrDeliveryEnded + if r := lastErr.Load(); r != nil { + reason = fmt.Errorf("%w: %w", ErrDeliveryEnded, *r) } - failed <- ErrDeliveryEnded + f.fail(fmt.Errorf("tenant %s: %w", id, reason)) }() - stop := func() { - stopped.Store(true) - cctx.Stop() + return nil +} + +// start begins delivery to deliver from every tenant's durable, and from +// each queue joined later, fetching about prefetch messages ahead across the +// tenants together (see share); watch reports a delivery that ends on its own +// through fail. The returned stop ends every delivery and stops joining new +// queues, without waiting. +func (f *fanIn) start(deliver func(jetstream.Msg), prefetch int, watch bool) (stop func(), err error) { + f.e.mu.Lock() + defer f.e.mu.Unlock() + f.deliver, f.prefetch, f.watch = deliver, prefetch, watch + stop = func() { + f.e.mu.Lock() + defer f.e.mu.Unlock() + f.stopped.Store(true) + for _, cctx := range f.running { + cctx.Stop() + } + f.e.unregister(f) + } + for _, id := range slices.Sorted(maps.Keys(f.handles)) { + if err := f.run(id); err != nil { + f.stopped.Store(true) + for _, cctx := range f.running { + cctx.Stop() + } + f.e.unregister(f) + return nil, fmt.Errorf("tenant %s: %w", id, err) + } + } + return stop, nil +} + +// workerConsumer is the Consumer CreateConsumer returns: a fanIn with the +// failed channel its contract promises. +type workerConsumer struct { + *fanIn + failed chan error +} + +func (c *workerConsumer) Consume(handler func(msg *Message), prefetch int) (func(), <-chan error, error) { + stop, err := c.start(func(m jetstream.Msg) { + handler(wrapMsg(c.ctx, m)) + }, prefetch, true) + if err != nil { + return nil, nil, fmt.Errorf("consume: %w", err) } - return stop, failed, nil + return stop, c.failed, nil } // stream resolves a stream handle by name. @@ -433,40 +817,71 @@ func (s *jsStream) consumerAckFloor(ctx context.Context, consumer string) (uint6 return info.AckFloor.Stream, nil } -// PurgeAcked purges the ingest stream below MIN(consumer's ack floor + 1, -// first sequence stored at or after olderThan) — see purgeAcked. -func (e *EmbeddedNATS) PurgeAcked(ctx context.Context, consumer string, olderThan time.Time) (bool, error) { - s, err := e.stream(ctx, ingestStream) - if err != nil { - return false, fmt.Errorf("get stream: %w", err) +// PurgeAcked purges each tenant's ingest stream below MIN(consumer's ack +// floor + 1, first sequence stored at or after the tenant's cutoff) — see +// purgeAcked. A tenant olderThan does not name is purged up to its ack floor. +// A failure on one tenant's stream is joined into the error and the sweep +// goes on to the next; a done ctx ends it. +func (e *EmbeddedNATS) PurgeAcked(ctx context.Context, consumer string, olderThan map[tenant.ID]time.Time) (bool, error) { + e.mu.Lock() + ids := e.ingestTenants() + e.mu.Unlock() + + now := time.Now() + var ( + errs []error + tenants int + ) + for _, id := range ids { + if err := ctx.Err(); err != nil { + errs = append(errs, err) + break + } + cutoff, ok := olderThan[id] + if !ok { + cutoff = now + } + s, err := e.stream(ctx, ingestStreamName(id)) + if err != nil { + errs = append(errs, fmt.Errorf("tenant %s: get stream: %w", id, err)) + continue + } + report, err := purgeAcked(ctx, s, consumer, cutoff) + if err != nil { + errs = append(errs, fmt.Errorf("tenant %s: %w", id, err)) + continue + } + // The sweep's own log lines: their detail is in sequences, which only + // this package speaks. Per tenant at Debug, since a sweep reaches + // every tenant each minute; the summary below is the Info line. + switch { + case report.purged: + tenants++ + slog.DebugContext(ctx, "sweeper: purged", + "tenant", id, + "purged_below_seq", report.target, + "ack_floor", report.ackFloor, + "gap_seq", report.gapSeq, + ) + case report.gapSeq == 0: + slog.DebugContext(ctx, "sweeper: all messages within gap window, skipping purge", "tenant", id) + } } - report, err := purgeAcked(ctx, s, consumer, olderThan) - if err != nil { - return false, err + if tenants > 0 { + slog.InfoContext(ctx, "sweeper: purged", "tenants", tenants) } - // The sweep's own log lines: their detail is in sequences, which only - // this package speaks. - switch { - case report.purged: - slog.InfoContext(ctx, "sweeper: purged", - "purged_below_seq", report.target, - "ack_floor", report.ackFloor, - "gap_seq", report.gapSeq, - ) - case report.gapSeq == 0: - slog.DebugContext(ctx, "sweeper: all messages within gap window, skipping purge") - } - return report.purged, nil -} - -// DeadLetterCounts reads the DLQ stream's per-subject counts and keys them by -// table across every tenant (see DeadLetterCounts.Tables). The table filter -// matches that table's unscoped subject under any tenant, so it is applied -// to the parsed topic rather than as a subject filter; a scoped topic counts -// under "table.scope". A subject written before the tenant led it counts -// under its table like any other (parseTopicKey). -func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, table string) (DeadLetterCounts, error) { - s, err := e.stream(ctx, dlqStream) + return tenants > 0, errors.Join(errs...) +} + +// DeadLetterCounts reads tenant id's dead-letter stream's per-subject counts +// and keys them by table. The table filter matches that table's unscoped +// subject, so it is applied to the parsed topic rather than as a subject +// filter; a scoped topic counts under "table.scope". +func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) { + if _, err := tenant.Parse(string(id)); err != nil { + return DeadLetterCounts{}, fmt.Errorf("tenant: %w", err) + } + s, err := e.stream(ctx, dlqStreamName(id)) if err != nil { if errors.Is(err, jetstream.ErrStreamNotFound) { return DeadLetterCounts{}, fmt.Errorf("%w: %w", ErrNoDeadLetterQueue, err) @@ -474,7 +889,7 @@ func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, table string) (Dead return DeadLetterCounts{}, fmt.Errorf("get dlq stream: %w", err) } - state, err := s.state(ctx, dlqAll) + state, err := s.state(ctx, tenantSubjects(dlqPrefix, id)) if err != nil { return DeadLetterCounts{}, fmt.Errorf("dlq stream info: %w", err) } @@ -495,21 +910,20 @@ func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, table string) (Dead return counts, nil } -// ReplaySince creates an ephemeral consumer on topic's ingest subject starting at -// since (DeliverByStartTime) and drains it to send until caught up. The -// consumer is ack-less and expires on its own once idle. Caught up is the -// client's no-messages or request-timeout answer to a pull; any other pull -// failure (a closed connection, a deleted consumer) is returned so the caller -// knows the replay ended short rather than empty. A done ctx ends the drain -// between pulls and returns ctx's error. A topic without a valid tenant is -// refused like a publish (see subject): the subject it names is exact, so -// events published before the tenant led the subject are not replayed. +// ReplaySince creates an ephemeral consumer on topic's ingest subject, in its +// tenant's queue, starting at since (DeliverByStartTime) and drains it to send +// until caught up. The consumer is ack-less and expires on its own once idle. +// Caught up is the client's no-messages or request-timeout answer to a pull; +// any other pull failure (a closed connection, a deleted consumer) is returned +// so the caller knows the replay ended short rather than empty. A done ctx +// ends the drain between pulls and returns ctx's error. A topic without a +// valid tenant is refused like a publish (see subject). func (e *EmbeddedNATS) ReplaySince(ctx context.Context, topic Topic, since time.Time, send func(data []byte) bool) error { subj, err := subject(ingestPrefix, topic) if err != nil { return err } - cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStream, jetstream.ConsumerConfig{ + cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStreamName(topic.Tenant), jetstream.ConsumerConfig{ FilterSubject: subj, DeliverPolicy: jetstream.DeliverByStartTimePolicy, OptStartTime: &since, diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index ab355df7..fe4fd45e 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -2,6 +2,9 @@ package mq import ( "context" + "fmt" + "os" + "path/filepath" "sync" "testing" "time" @@ -13,16 +16,62 @@ import ( "github.com/stretchr/testify/require" ) -// newTestEmbedded spins up an EmbeddedNATS with a temporary store directory -// that is cleaned up by the test framework. -func newTestEmbedded(t *testing.T) *EmbeddedNATS { +// testBudget is the byte budget newTestEmbedded opens each queue at. +const testBudget = 64 << 20 + +// openEmbedded starts an EmbeddedNATS over dir, closed by the test framework. +func openEmbedded(t *testing.T, dir string) *EmbeddedNATS { t.Helper() - e, err := NewEmbedded(t.TempDir(), 64<<20) + e, err := NewEmbedded(dir) require.NoError(t, err) t.Cleanup(func() { _ = e.Close() }) return e } +// newTestEmbedded spins up an EmbeddedNATS over a temporary store directory +// with a queue open for each of tenants — tenant.Default when none is named — +// at testBudget. +func newTestEmbedded(t *testing.T, tenants ...tenant.ID) *EmbeddedNATS { + t.Helper() + e := openEmbedded(t, t.TempDir()) + if len(tenants) == 0 { + tenants = []tenant.ID{tenant.Default} + } + for _, id := range tenants { + require.NoError(t, e.SetMaxBytes(t.Context(), id, testBudget)) + } + return e +} + +// streamConfig is the stored config of the named stream. +func streamConfig(t *testing.T, e *EmbeddedNATS, name string) jetstream.StreamConfig { + t.Helper() + s, err := e.js.Stream(t.Context(), name) + require.NoError(t, err) + return s.CachedInfo().Config +} + +// ackAll consumes every message delivered to consumer on the ingest queue, +// acknowledging each, until n have been acked. +func ackAll(t *testing.T, e *EmbeddedNATS, consumer string, n int) { + t.Helper() + ctx := t.Context() + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: consumer, MaxAckPending: 100}) + require.NoError(t, err) + acked := make(chan error, n) + stop, _, err := cons.Consume(func(msg *Message) { acked <- msg.DoubleAck(ctx) }, 10) + require.NoError(t, err) + t.Cleanup(stop) + for range n { + select { + case err := <-acked: + require.NoError(t, err) + case <-time.After(5 * time.Second): + t.Fatal("timed out waiting for acks") + } + } +} + func TestEmbeddedNATS_PublishSubscribe(t *testing.T) { // No t.Parallel(): each embedded server uses DontListen+InProcessServer, // but starting several in parallel still slows tests unnecessarily. @@ -83,7 +132,7 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // Read the stored message back raw: the option headers are on the wire // exactly as set, exact-key, with Add appending rather than replacing. - s, err := e.js.Stream(ctx, ingestStream) + s, err := e.js.Stream(ctx, "INGEST_0") require.NoError(t, err) raw, err := s.GetLastMsgForSubject(ctx, "ingest.0.hdr") require.NoError(t, err) @@ -92,24 +141,30 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { assert.Equal(t, []byte("x"), raw.Data) } -func TestNewEmbedded_CreatesBothStreams(t *testing.T) { - e := newTestEmbedded(t) - ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) - defer cancel() - - assert.Equal(t, int64(64<<20), e.MaxBytes()) - - ingest, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20), ingest.CachedInfo().Config.MaxBytes) - - // The DLQ stream is always present, at a tenth of the budget. - dlq, err := e.js.Stream(ctx, dlqStream) - require.NoError(t, err) - cfg := dlq.CachedInfo().Config - assert.Equal(t, []string{"dlq.>"}, cfg.Subjects) - assert.Equal(t, int64(64<<20)/10, cfg.MaxBytes) - assert.Equal(t, jetstream.DiscardOld, cfg.Discard) +// A tenant's first budget opens its queue: an ingest stream holding its +// subjects alone at the budget, refusing when full, and a dead-letter stream +// at a tenth of it, dropping its oldest when full. No other tenant gets one. +func TestEmbeddedNATS_SetMaxBytes_OpensTheTenantsQueue(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + assert.Zero(t, e.MaxBytes("acme"), "no budget applied yet") + + require.NoError(t, e.SetMaxBytes(t.Context(), "acme", testBudget)) + assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) + + ingest := streamConfig(t, e, "INGEST_acme") + assert.Equal(t, []string{"ingest.acme.>"}, ingest.Subjects) + assert.Equal(t, int64(testBudget), ingest.MaxBytes) + assert.Equal(t, jetstream.DiscardNew, ingest.Discard) + dlq := streamConfig(t, e, "DLQ_acme") + assert.Equal(t, []string{"dlq.acme.>"}, dlq.Subjects) + assert.Equal(t, int64(testBudget)/10, dlq.MaxBytes) + assert.Equal(t, jetstream.DiscardOld, dlq.Discard) + + assert.Zero(t, e.MaxBytes("globex")) + _, err := e.js.Stream(t.Context(), "INGEST_globex") + require.ErrorIs(t, err, jetstream.ErrStreamNotFound, "another tenant's queue opens with its own budget") + + require.Error(t, e.SetMaxBytes(t.Context(), "a.b", testBudget), "a tenant outside the grammar has no queue") } func TestEmbeddedNATS_StreamHandle(t *testing.T) { @@ -120,7 +175,7 @@ func TestEmbeddedNATS_StreamHandle(t *testing.T) { _, err := e.stream(ctx, "NO_SUCH_STREAM") require.Error(t, err, "an unknown stream is an error, not a nil handle") - s, err := e.stream(ctx, ingestStream) + s, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) empty, err := s.state(ctx, "") @@ -196,11 +251,11 @@ func TestEmbeddedNATS_StreamHandle(t *testing.T) { } // TestEmbeddedNATS_CreateConsumer_Config pins the ConsumerConfig → broker -// mapping: AckWait (redelivery timing) and MaxAckPending (ingest backpressure) -// are checkable nowhere else, and a dropped field would compile and pass -// every delivery test. +// mapping on every tenant's queue: AckWait (redelivery timing) and +// MaxAckPending (ingest backpressure, per tenant) are checkable nowhere else, +// and a dropped field would compile and pass every delivery test. func TestEmbeddedNATS_CreateConsumer_Config(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() @@ -211,17 +266,16 @@ func TestEmbeddedNATS_CreateConsumer_Config(t *testing.T) { }) require.NoError(t, err) - s, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - cons, err := s.Consumer(ctx, "cfg") - require.NoError(t, err) - info, err := cons.Info(ctx) - require.NoError(t, err) - assert.Equal(t, "cfg", info.Config.Durable) - assert.Equal(t, "ingest.>", info.Config.FilterSubject, "the consumer sees every topic") - assert.Equal(t, jetstream.AckExplicitPolicy, info.Config.AckPolicy) - assert.Equal(t, 42*time.Second, info.Config.AckWait) - assert.Equal(t, 123, info.Config.MaxAckPending) + for _, stream := range []string{"INGEST_acme", "INGEST_globex"} { + cons, err := e.js.Consumer(ctx, stream, "cfg") + require.NoError(t, err, stream) + cfg := cons.CachedInfo().Config + assert.Equal(t, "cfg", cfg.Durable) + assert.Empty(t, cfg.FilterSubject, "%s: the durable sees the whole of its tenant's stream", stream) + assert.Equal(t, jetstream.AckExplicitPolicy, cfg.AckPolicy) + assert.Equal(t, 42*time.Second, cfg.AckWait) + assert.Equal(t, 123, cfg.MaxAckPending) + } } func TestEmbeddedNATS_ReplaySince(t *testing.T) { @@ -262,7 +316,7 @@ func TestEmbeddedNATS_ReplaySince(t *testing.T) { func TestEmbeddedNATS_DefaultLogger(t *testing.T) { // NewEmbedded without a logger should not panic — it falls back to the // default slog logger. - e, err := NewEmbedded(t.TempDir(), 64<<20) + e, err := NewEmbedded(t.TempDir()) require.NoError(t, err) t.Cleanup(func() { _ = e.Close() }) } @@ -301,29 +355,32 @@ func TestSlogNATSLogger_Levels(t *testing.T) { } func TestEmbeddedNATS_SetMaxBytes(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - require.NoError(t, e.SetMaxBytes(ctx, 128<<20)) - assert.Equal(t, int64(128<<20), e.MaxBytes()) + require.NoError(t, e.SetMaxBytes(ctx, "acme", 128<<20)) + assert.Equal(t, int64(128<<20), e.MaxBytes("acme")) - ingest, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - assert.Equal(t, int64(128<<20), ingest.CachedInfo().Config.MaxBytes) + ingest := streamConfig(t, e, "INGEST_acme") + assert.Equal(t, int64(128<<20), ingest.MaxBytes) // Everything but the limit is preserved. - assert.Equal(t, []string{"ingest.>"}, ingest.CachedInfo().Config.Subjects) - assert.Equal(t, jetstream.DiscardNew, ingest.CachedInfo().Config.Discard) + assert.Equal(t, []string{"ingest.acme.>"}, ingest.Subjects) + assert.Equal(t, jetstream.DiscardNew, ingest.Discard) - // The DLQ stream follows at a tenth of the budget. - dlq, err := e.js.Stream(ctx, dlqStream) - require.NoError(t, err) - assert.Equal(t, int64(128<<20)/10, dlq.CachedInfo().Config.MaxBytes) - assert.Equal(t, jetstream.DiscardOld, dlq.CachedInfo().Config.Discard) + // The dead-letter stream follows at a tenth of the budget. + dlq := streamConfig(t, e, "DLQ_acme") + assert.Equal(t, int64(128<<20)/10, dlq.MaxBytes) + assert.Equal(t, jetstream.DiscardOld, dlq.Discard) + + // No other tenant's queue moves. + assert.Equal(t, int64(testBudget), e.MaxBytes("globex")) + assert.Equal(t, int64(testBudget), streamConfig(t, e, "INGEST_globex").MaxBytes) + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_globex").MaxBytes) // The budget already in effect is a no-op, not an error. - require.NoError(t, e.SetMaxBytes(ctx, 128<<20)) - assert.Equal(t, int64(128<<20), e.MaxBytes()) + require.NoError(t, e.SetMaxBytes(ctx, "acme", 128<<20)) + assert.Equal(t, int64(128<<20), e.MaxBytes("acme")) } func TestEmbeddedNATS_SetMaxBytes_DLQFailureRollsBackIngest(t *testing.T) { @@ -331,27 +388,23 @@ func TestEmbeddedNATS_SetMaxBytes_DLQFailureRollsBackIngest(t *testing.T) { ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - // Put the DLQ stream where the update can't follow: JetStream refuses to - // change a live stream's retention policy, so recreating it as a work - // queue makes the DLQ resize fail after the ingest resize has already - // succeeded. - require.NoError(t, e.js.DeleteStream(ctx, dlqStream)) + // Put the dead-letter stream where the update can't follow: JetStream + // refuses to change a live stream's retention policy, so recreating it as + // a work queue makes the dead-letter resize fail after the ingest resize + // has already succeeded. + require.NoError(t, e.js.DeleteStream(ctx, "DLQ_0")) _, err := e.js.CreateStream(ctx, jetstream.StreamConfig{ - Name: dlqStream, Subjects: []string{dlqAll}, Retention: jetstream.WorkQueuePolicy, MaxBytes: (64 << 20) / 10, + Name: "DLQ_0", Subjects: []string{"dlq.0.>"}, Retention: jetstream.WorkQueuePolicy, MaxBytes: testBudget / 10, }) require.NoError(t, err) - err = e.SetMaxBytes(ctx, 128<<20) + err = e.SetMaxBytes(ctx, tenant.Default, 128<<20) require.Error(t, err) assert.Contains(t, err.Error(), "ingest stream restored to the previous limit") - assert.Equal(t, int64(64<<20), e.MaxBytes(), "the budget in effect is unchanged, so the next call retries both") + assert.Equal(t, int64(testBudget), e.MaxBytes(tenant.Default), "the budget in effect is unchanged, so the next call retries both") - ingest, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20), ingest.CachedInfo().Config.MaxBytes, "the ingest resize is undone so the pair stays at the previous limit") - dlq, err := e.js.Stream(ctx, dlqStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20)/10, dlq.CachedInfo().Config.MaxBytes) + assert.Equal(t, int64(testBudget), streamConfig(t, e, "INGEST_0").MaxBytes, "the ingest resize is undone so the pair stays at the previous limit") + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_0").MaxBytes) } func TestEmbeddedNATS_SetMaxBytes_IngestFailureChangesNothing(t *testing.T) { @@ -359,13 +412,86 @@ func TestEmbeddedNATS_SetMaxBytes_IngestFailureChangesNothing(t *testing.T) { ctx, cancel := context.WithCancel(t.Context()) cancel() // a stop caught mid-reload: the first JetStream call gives up - err := e.SetMaxBytes(ctx, 128<<20) + err := e.SetMaxBytes(ctx, tenant.Default, 128<<20) require.ErrorIs(t, err, context.Canceled) - assert.Equal(t, int64(64<<20), e.MaxBytes()) + assert.Equal(t, int64(testBudget), e.MaxBytes(tenant.Default)) - dlq, err := e.js.Stream(t.Context(), dlqStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20)/10, dlq.CachedInfo().Config.MaxBytes, "the dlq is not touched when the ingest resize fails") + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_0").MaxBytes, "the dead-letter stream is not touched when the ingest resize fails") +} + +// A tenant whose queue JetStream will not open — here, a file where its +// dead-letter stream's store would go — is refused on its own: SetMaxBytes +// errors and applies no budget, and a publish is refused as a full queue, +// while every other tenant's queue opens after it (which a store limit at +// the very top of the int64 range would refuse: see NewEmbedded). Once the +// cause is gone, a publish opens the queue at the budget last asked for it. +func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { + dir := t.TempDir() + // The dead-letter stream is the first of the pair to open. A failed open + // removes what was in the way, so the obstacle is put back before each + // attempt meant to fail. + block := filepath.Join(dir, "jetstream", "$G", "streams", dlqStreamName("acme")) + obstruct := func() { + t.Helper() + require.NoError(t, os.MkdirAll(filepath.Dir(block), 0o750)) + require.NoError(t, os.WriteFile(block, nil, 0o600)) + } + obstruct() + e := openEmbedded(t, dir) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + require.Error(t, e.SetMaxBytes(ctx, "acme", testBudget)) + assert.Zero(t, e.MaxBytes("acme"), "no budget applied") + + require.NoError(t, e.SetMaxBytes(ctx, "globex", testBudget), "one tenant's failed open costs the next nothing") + require.NoError(t, e.Publish(ctx, Topic{Tenant: "globex", Table: "t"}, []byte("x"))) + + obstruct() + err := e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x")) + require.ErrorIs(t, err, ErrQueueFull, "the tenant's queue takes nothing; a retry is the answer") + + if err := os.Remove(block); err != nil { + require.ErrorIs(t, err, os.ErrNotExist) + } + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) +} + +// A budget that shrinks a tenant's dead-letter stream below what it holds +// would have DiscardOld delete the oldest parked rows to fit (#532), so the +// stream keeps what it holds, capped at that, and every row survives. +func TestEmbeddedNATS_SetMaxBytes_NeverShrinksTheDeadLetterQueueBelowWhatItHolds(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + require.NoError(t, e.SetMaxBytes(ctx, "acme", 10<<20)) + + payload := make([]byte, 1<<10) + for range 200 { + msg := NewMessage(ctx, Topic{Tenant: "acme", Table: "t"}, payload, time.Now(), nil, nil, nil) + require.NoError(t, e.DeadLetter(ctx, msg)) + } + dlqState := func() jetstream.StreamState { + t.Helper() + s, err := e.js.Stream(ctx, "DLQ_acme") + require.NoError(t, err) + return s.CachedInfo().State + } + held := dlqState().Bytes + require.Greater(t, held, uint64(100<<10), "the rows take more than a tenth of the budget below") + + // Shrunk to a 1 MB budget: a tenth of it is less than the stream holds. + require.NoError(t, e.SetMaxBytes(ctx, "acme", 1<<20)) + assert.Equal(t, int64(1<<20), e.MaxBytes("acme"), "the budget applies") + assert.Equal(t, int64(1<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) + assert.Equal(t, held, uint64(streamConfig(t, e, "DLQ_acme").MaxBytes), "capped at what it holds, not at a tenth") //nolint:gosec // G115: a stream cap is never negative + assert.Equal(t, uint64(200), dlqState().Msgs, "no parked row is deleted") + + // A budget whose tenth covers what it holds applies as usual. + require.NoError(t, e.SetMaxBytes(ctx, "acme", 4<<20)) + assert.Equal(t, int64(4<<20)/10, streamConfig(t, e, "DLQ_acme").MaxBytes) + assert.Equal(t, uint64(200), dlqState().Msgs) } func TestEmbeddedNATS_ReplaySince_PullFailureIsAnError(t *testing.T) { @@ -410,11 +536,11 @@ func TestEmbeddedNATS_ReplaySince_StopsWhenContextIsDone(t *testing.T) { } func TestEmbeddedNATS_DeadLetter(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, tenant.Default, "acme") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - empty, err := e.DeadLetterCounts(ctx, "") + empty, err := e.DeadLetterCounts(ctx, tenant.Default, "") require.NoError(t, err) assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{}}, empty) @@ -426,52 +552,58 @@ func TestEmbeddedNATS_DeadLetter(t *testing.T) { park(Topic{Tenant: tenant.Default, Table: "default.orders"}, "o1") park(Topic{Tenant: tenant.Default, Table: "default.orders"}, "o2") park(Topic{Tenant: tenant.Default, Table: "users"}, "u1") - // Another tenant's table of the same name counts with it: one queue, one - // count, until the queue is per tenant. So does a subject parked before - // the tenant led it — the queue is never drained, so those stay. + // Another tenant's table of the same name is its own queue and its own + // count. park(Topic{Tenant: "acme", Table: "users"}, "acme-u1") - _, err = e.js.Publish(ctx, "dlq.users", []byte("pre-tenant")) - require.NoError(t, err) - // Parked under the same topic on the DLQ stream, headers intact, and - // nothing lands on the ingest stream. - dlq, err := e.js.Stream(ctx, dlqStream) + // Parked under the same topic on the tenant's dead-letter stream, headers + // intact, and nothing lands on the ingest stream. + dlq, err := e.js.Stream(ctx, "DLQ_0") require.NoError(t, err) raw, err := dlq.GetLastMsgForSubject(ctx, "dlq.0.default%2Eorders") require.NoError(t, err) assert.Equal(t, []byte("o2"), raw.Data) assert.Equal(t, "boom", raw.Header.Get("X-DLQ-Error")) - ingest, err := e.stream(ctx, ingestStream) + ingest, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) st, err := ingest.state(ctx, "") require.NoError(t, err) assert.Zero(t, st.Msgs) - all, err := e.DeadLetterCounts(ctx, "") + all, err := e.DeadLetterCounts(ctx, tenant.Default, "") require.NoError(t, err) - assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2, "users": 3}, Total: 5}, all, "table names come back decoded, summed across tenants") + assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2, "users": 1}, Total: 3}, all, "table names come back decoded, the tenant's own alone") - one, err := e.DeadLetterCounts(ctx, "default.orders") + one, err := e.DeadLetterCounts(ctx, tenant.Default, "default.orders") require.NoError(t, err) - assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2}, Total: 5}, one, "Total is every parked message, filter or not") + assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2}, Total: 3}, one, "Total is every parked message of the tenant, filter or not") - users, err := e.DeadLetterCounts(ctx, "users") + acme, err := e.DeadLetterCounts(ctx, "acme", "") require.NoError(t, err) - assert.Equal(t, map[string]uint64{"users": 3}, users.Tables, "the filter is by table under any tenant, the pre-tenant subject included") + assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"users": 1}, Total: 1}, acme) - none, err := e.DeadLetterCounts(ctx, "never_failed") + none, err := e.DeadLetterCounts(ctx, tenant.Default, "never_failed") require.NoError(t, err) assert.Empty(t, none.Tables) } +// A tenant with no queue — one never given a budget on this data directory — +// has nothing parked, which is not the same as a failed read. func TestEmbeddedNATS_DeadLetterCounts_NoQueue(t *testing.T) { e := newTestEmbedded(t) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - require.NoError(t, e.js.DeleteStream(ctx, dlqStream)) - _, err := e.DeadLetterCounts(ctx, "") + _, err := e.DeadLetterCounts(ctx, "globex", "") + require.ErrorIs(t, err, ErrNoDeadLetterQueue) + + require.NoError(t, e.js.DeleteStream(ctx, "DLQ_0")) + _, err = e.DeadLetterCounts(ctx, tenant.Default, "") require.ErrorIs(t, err, ErrNoDeadLetterQueue) + + _, err = e.DeadLetterCounts(ctx, "a.b", "") + require.Error(t, err, "an id outside the grammar names no stream") + assert.NotErrorIs(t, err, ErrNoDeadLetterQueue) } func TestEmbeddedNATS_DeadLetterCounts_BrokerFailureIsNotAnEmptyQueue(t *testing.T) { @@ -482,19 +614,19 @@ func TestEmbeddedNATS_DeadLetterCounts_BrokerFailureIsNotAnEmptyQueue(t *testing // A lookup that fails for any reason other than "no such stream" must not // read as an empty queue. e.conn.Close() - _, err := e.DeadLetterCounts(ctx, "") + _, err := e.DeadLetterCounts(ctx, tenant.Default, "") require.Error(t, err) assert.NotErrorIs(t, err, ErrNoDeadLetterQueue) } func TestEmbeddedNATS_DeadLetter_IsAPrefixSwap(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "a") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() // A subject this package would never write (four tokens) still parks - // under the very same tail: nothing on the dead-letter path decodes or - // re-encodes it. + // under the very same tail, in the queue of the tenant its first token + // names: nothing on the dead-letter path decodes or re-encodes it. _, err := e.js.Publish(ctx, "ingest.a.b.c.d", []byte("foreign")) require.NoError(t, err) @@ -513,29 +645,72 @@ func TestEmbeddedNATS_DeadLetter_IsAPrefixSwap(t *testing.T) { assert.Equal(t, Topic{Table: "a.b.c.d"}, msg.Topic(), "a foreign tail is the table of no tenant") require.NoError(t, e.DeadLetter(ctx, msg)) - dlq, err := e.js.Stream(ctx, dlqStream) + dlq, err := e.js.Stream(ctx, "DLQ_a") require.NoError(t, err) raw, err := dlq.GetLastMsgForSubject(ctx, "dlq.a.b.c.d") require.NoError(t, err) assert.Equal(t, []byte("foreign"), raw.Data) } -func TestEmbeddedNATS_Publish_QueueFull(t *testing.T) { - e, err := NewEmbedded(t.TempDir(), 4<<10) +// A dead-letter stream that has gone missing is opened again with its +// tenant's queue, at a tenth of the budget last asked for it, rather than +// leaving the row to be redelivered. +func TestEmbeddedNATS_DeadLetter_ReopensAMissingQueue(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + require.NoError(t, e.js.DeleteStream(ctx, "DLQ_acme")) + require.NoError(t, e.DeadLetter(ctx, NewMessage(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"), time.Now(), nil, nil, nil))) + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_acme").MaxBytes) + counts, err := e.DeadLetterCounts(ctx, "acme", "") require.NoError(t, err) - t.Cleanup(func() { _ = e.Close() }) + assert.Equal(t, uint64(1), counts.Total) +} + +func TestEmbeddedNATS_Publish_QueueFull(t *testing.T) { + e := openEmbedded(t, t.TempDir()) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() + require.NoError(t, e.SetMaxBytes(ctx, "acme", 4<<10)) + require.NoError(t, e.SetMaxBytes(ctx, "globex", 4<<10)) // DiscardNew refuses the publish that would pass the byte budget; that is // the backpressure signal, named so callers need not read broker errors. payload := make([]byte, 1<<10) + var err error for range 8 { - if err = e.Publish(ctx, Topic{Tenant: tenant.Default, Table: "full"}, payload); err != nil { + if err = e.Publish(ctx, Topic{Tenant: "acme", Table: "full"}, payload); err != nil { break } } require.ErrorIs(t, err, ErrQueueFull) + + // Only the tenant at its budget is refused: the next one has a budget of + // its own. + require.NoError(t, e.Publish(ctx, Topic{Tenant: "globex", Table: "full"}, payload)) +} + +// A tenant's queue opens at the budget last asked for it when a publish finds +// it missing, and a tenant never given a budget has no queue to publish to: +// that is refused as a full queue, and nothing is opened for it. +func TestEmbeddedNATS_Publish_OpensTheQueueAtTheLastBudget(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + for _, name := range []string{"INGEST_acme", "DLQ_acme"} { + require.NoError(t, e.js.DeleteStream(ctx, name)) + } + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(testBudget), streamConfig(t, e, "INGEST_acme").MaxBytes) + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_acme").MaxBytes) + + err := e.Publish(ctx, Topic{Tenant: "globex", Table: "t"}, []byte("x")) + require.ErrorIs(t, err, ErrQueueFull) + assert.Contains(t, err.Error(), "globex") + _, err = e.js.Stream(ctx, "INGEST_globex") + require.ErrorIs(t, err, jetstream.ErrStreamNotFound) } func TestEmbeddedNATS_PurgeAcked(t *testing.T) { @@ -544,7 +719,7 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { defer cancel() // No consumer yet: the sentinel the sweeper keys its "not yet" warning on. - _, err := e.PurgeAcked(ctx, "buffer", time.Now()) + _, err := e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{tenant.Default: time.Now()}) require.ErrorIs(t, err, ErrConsumerNotFound) for i := range 4 { @@ -570,7 +745,7 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { t.Fatal("timed out waiting for acks") } } - s, err := e.stream(ctx, ingestStream) + s, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) require.Eventually(t, func() bool { floor, err := s.consumerAckFloor(ctx, "buffer") @@ -578,12 +753,12 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { }, 5*time.Second, 20*time.Millisecond) // Everything is acked-or-not but nothing is old enough: keep it all. - purged, err := e.PurgeAcked(ctx, "buffer", time.Now().Add(-time.Hour)) + purged, err := e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{tenant.Default: time.Now().Add(-time.Hour)}) require.NoError(t, err) assert.False(t, purged) // Everything is old enough: only the acked two go. - purged, err = e.PurgeAcked(ctx, "buffer", time.Now().Add(time.Hour)) + purged, err = e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{tenant.Default: time.Now().Add(time.Hour)}) require.NoError(t, err) assert.True(t, purged) st, err := s.state(ctx, "") @@ -592,8 +767,147 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { assert.Equal(t, uint64(2), st.Msgs) } +// Each tenant's queue is purged at its own cutoff and below its own ack +// floor: a tenant keeping an hour of history keeps it while the next one's +// goes, and a tenant the cutoffs do not name — one no longer served — keeps +// no history at all. +func TestEmbeddedNATS_PurgeAcked_EachTenantAtItsOwnCutoff(t *testing.T) { + e := newTestEmbedded(t, "acme", "globex", "initech") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + for _, id := range []tenant.ID{"acme", "globex", "initech"} { + for i := range 2 { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "p"}, []byte{byte(i)})) + } + } + ackAll(t, e, "buffer", 6) + for _, id := range []tenant.ID{"acme", "globex", "initech"} { + s, err := e.stream(ctx, ingestStreamName(id)) + require.NoError(t, err) + require.Eventually(t, func() bool { + floor, err := s.consumerAckFloor(ctx, "buffer") + return err == nil && floor == 2 + }, 5*time.Second, 20*time.Millisecond, id) + } + + purged, err := e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{ + "acme": time.Now().Add(-time.Hour), // an hour of history: all of it inside the window + "globex": time.Now().Add(time.Hour), // everything older than the cutoff + }) + require.NoError(t, err) + assert.True(t, purged) + msgs := func(id tenant.ID) uint64 { + s, err := e.stream(ctx, ingestStreamName(id)) + require.NoError(t, err) + st, err := s.state(ctx, "") + require.NoError(t, err) + return st.Msgs + } + assert.Equal(t, uint64(2), msgs("acme"), "kept for its own window") + assert.Zero(t, msgs("globex"), "past its own window") + assert.Zero(t, msgs("initech"), "a tenant the cutoffs do not name keeps nothing it has acknowledged") +} + +// The isolation per-tenant queues buy: a tenant at MaxAckPending, or one +// whose handler is stuck, holds back its own delivery and no other tenant's — +// each tenant's messages arrive on a delivery of their own, in order. +func TestEmbeddedNATS_Consume_OneTenantsBacklogDoesNotHoldAnother(t *testing.T) { + e := newTestEmbedded(t, "acme", "globex", "initech") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 2}) + require.NoError(t, err) + release := make(chan struct{}) + var mu sync.Mutex + delivered := map[tenant.ID][]byte{} + stop, _, err := cons.Consume(func(msg *Message) { + id := msg.Topic().Tenant + mu.Lock() + delivered[id] = append(delivered[id], msg.Data[0]) + mu.Unlock() + if id == "initech" { + <-release // never returns until the test ends + } + if id == "globex" { + _ = msg.Ack() // acme never acks: its delivery stops at MaxAckPending + } + }, 12) + require.NoError(t, err) + t.Cleanup(func() { + close(release) + stop() + }) + + for i := range 5 { + for _, id := range []tenant.ID{"acme", "globex", "initech"} { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "t"}, []byte{byte(i)})) + } + } + counts := func() (acme, globex, initech int) { + mu.Lock() + defer mu.Unlock() + return len(delivered["acme"]), len(delivered["globex"]), len(delivered["initech"]) + } + require.Eventually(t, func() bool { + acme, globex, initech := counts() + return acme == 2 && globex == 5 && initech == 1 + }, 5*time.Second, 20*time.Millisecond, "globex is delivered in full while acme waits on its acks and initech on its handler") + time.Sleep(200 * time.Millisecond) + acme, globex, initech := counts() + assert.Equal(t, 2, acme, "no more than MaxAckPending unacked, for acme alone") + assert.Equal(t, 5, globex) + assert.Equal(t, 1, initech, "a stuck handler holds back its own tenant alone") + mu.Lock() + defer mu.Unlock() + assert.Equal(t, []byte{0, 1, 2, 3, 4}, delivered["globex"], "in the order published") +} + +// A tenant's queue opened after the consumer started is joined to it: both +// consumer paths deliver its events as they do the queues that were there +// first, whether those were opened in this process or found on disk. +func TestEmbeddedNATS_ConsumersJoinQueuesOpenedLater(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + worker := make(chan Topic, 4) + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 10}) + require.NoError(t, err) + stop, _, err := cons.Consume(func(msg *Message) { + _ = msg.Ack() + worker <- msg.Topic() + }, 4) + require.NoError(t, err) + t.Cleanup(stop) + hub := make(chan Topic, 4) + require.NoError(t, e.Subscribe(ctx, "hub-bridge", func(msg *Message) error { + _ = msg.Ack() + hub <- msg.Topic() + return nil + })) + + require.NoError(t, e.SetMaxBytes(ctx, "globex", testBudget)) + for _, id := range []tenant.ID{"acme", "globex"} { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "t"}, []byte("x"))) + } + for name, got := range map[string]chan Topic{"worker": worker, "hub": hub} { + var topics []Topic + for range 2 { + select { + case topic := <-got: + topics = append(topics, topic) + case <-time.After(5 * time.Second): + t.Fatalf("%s: timed out; delivered %v", name, topics) + } + } + assert.ElementsMatch(t, []Topic{{Tenant: "acme", Table: "t"}, {Tenant: "globex", Table: "t"}}, topics, name) + } +} + func TestEmbeddedNATS_Consume_ReportsDeliveryEndingOnItsOwn(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second) defer cancel() @@ -609,22 +923,23 @@ func TestEmbeddedNATS_Consume_ReportsDeliveryEndingOnItsOwn(t *testing.T) { case <-time.After(200 * time.Millisecond): } - // Deleting the durable underneath a running Consume is terminal: the - // client stops the subscription on its own, and no message will ever say - // so. It must reach the caller. - require.NoError(t, e.js.DeleteConsumer(ctx, ingestStream, "doomed")) + // Deleting one tenant's durable underneath a running Consume is terminal + // for that tenant: the client stops the subscription on its own, and no + // message will ever say so. It must reach the caller. + require.NoError(t, e.js.DeleteConsumer(ctx, "INGEST_globex", "doomed")) select { case err := <-failed: require.ErrorIs(t, err, ErrDeliveryEnded) require.ErrorIs(t, err, jetstream.ErrConsumerDeleted, "the broker's reason is kept") + assert.Contains(t, err.Error(), "globex", "the tenant is named") case <-ctx.Done(): t.Fatal("delivery ended underneath the consumer and nothing was reported") } } func TestEmbeddedNATS_Consume_StopIsNotAFailure(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() @@ -639,6 +954,29 @@ func TestEmbeddedNATS_Consume_StopIsNotAFailure(t *testing.T) { t.Fatalf("our own stop was reported as a failure: %v", err) case <-time.After(time.Second): } + // Nor is a queue opened after the stop joined to it. + require.NoError(t, e.SetMaxBytes(ctx, "initech", testBudget)) + _, err = e.js.Consumer(ctx, "INGEST_initech", "stopped") + require.ErrorIs(t, err, jetstream.ErrConsumerNotFound) +} + +// The fetch-ahead asked for is shared by the tenants' queues, at least one +// each, so the rows held client-side stay about what the caller asked for +// however many tenants there are. +func TestFanIn_SharesThePrefetch(t *testing.T) { + t.Parallel() + handles := func(n int) map[tenant.ID]jetstream.Consumer { + m := map[tenant.ID]jetstream.Consumer{} + for i := range n { + m[tenant.ID(fmt.Sprint(i))] = nil + } + return m + } + assert.Equal(t, 500, (&fanIn{prefetch: 500, handles: handles(1)}).share()) + assert.Equal(t, 250, (&fanIn{prefetch: 500, handles: handles(2)}).share()) + assert.Equal(t, 1, (&fanIn{prefetch: 500, handles: handles(1000)}).share(), "at least one per tenant") + assert.Equal(t, 500, (&fanIn{prefetch: 500}).share(), "no tenant yet") + assert.Zero(t, (&fanIn{handles: handles(3)}).share(), "0 leaves the client default") } // Nothing lands on the default tenant by omission (#583): the tenant is a @@ -652,7 +990,7 @@ func TestEmbeddedNATS_Publish_RefusesATopicWithoutATenant(t *testing.T) { require.Error(t, e.Publish(ctx, topic, []byte("x")), "%+v", topic) require.Error(t, e.ReplaySince(ctx, topic, time.Time{}, func([]byte) bool { return true }), "%+v", topic) } - s, err := e.stream(ctx, ingestStream) + s, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) st, err := s.state(ctx, "") require.NoError(t, err) @@ -661,7 +999,7 @@ func TestEmbeddedNATS_Publish_RefusesATopicWithoutATenant(t *testing.T) { // Two tenants, one table name: a replay of one never carries the other's rows. func TestEmbeddedNATS_ReplaySince_IsPerTenant(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() @@ -677,34 +1015,69 @@ func TestEmbeddedNATS_ReplaySince_IsPerTenant(t *testing.T) { assert.Equal(t, []string{"acme1", "acme2"}, got) } -// A message published before the tenant led the subject (#583 story 5) is -// still delivered after the upgrade — the durable consumers filter ingest.> -// — and reads as the default tenant's, so it inserts, streams and parks as -// it did. -func TestEmbeddedNATS_PreTenantSubjectsStillDeliver(t *testing.T) { - e := newTestEmbedded(t) +// A boot over a directory an earlier build wrote deletes the pair of streams +// it kept for every tenant together: their subjects overlap every tenant's, +// so no tenant's queue could open beside them. +func TestNewEmbedded_DeletesTheStreamsAnEarlierBuildShared(t *testing.T) { + dir := t.TempDir() + old, err := NewEmbedded(dir) + require.NoError(t, err) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - - _, err := e.js.Publish(ctx, "ingest.events", []byte("old")) + for name, subj := range map[string]string{legacyIngestStream: "ingest.>", legacyDLQStream: "dlq.>"} { + _, err := old.js.CreateStream(ctx, jetstream.StreamConfig{Name: name, Subjects: []string{subj}}) + require.NoError(t, err) + } + _, err = old.js.Publish(ctx, "ingest.events", []byte("pre-tenant")) require.NoError(t, err) + require.NoError(t, old.Close()) - got := make(chan *Message, 1) - require.NoError(t, e.Subscribe(ctx, "upgrade", func(msg *Message) error { - got <- msg - return nil - })) - var msg *Message - select { - case msg = <-got: - case <-time.After(5 * time.Second): - t.Fatal("timed out waiting for delivery") + e := openEmbedded(t, dir) + for _, name := range []string{legacyIngestStream, legacyDLQStream} { + _, err := e.js.Stream(ctx, name) + require.ErrorIs(t, err, jetstream.ErrStreamNotFound, name) } - assert.Equal(t, Topic{Tenant: tenant.Default, Table: "events"}, msg.Topic()) - require.NoError(t, e.DeadLetter(ctx, msg)) - dlq, err := e.js.Stream(ctx, dlqStream) + require.NoError(t, e.SetMaxBytes(ctx, tenant.Default, testBudget)) + require.NoError(t, e.Publish(ctx, Topic{Tenant: tenant.Default, Table: "events"}, []byte("x"))) +} + +// A boot takes stock of the queues on disk: each keeps the budget it last +// had, and a consumer created afterwards is held on every one of them — a +// tenant no longer served, which is never given a budget again, included — +// so what such a tenant had queued still reaches the worker. +func TestNewEmbedded_TakesStockOfTheQueuesOnDisk(t *testing.T) { + dir := t.TempDir() + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + first, err := NewEmbedded(dir) + require.NoError(t, err) + require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) + for i := range 2 { + require.NoError(t, first.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte{byte(i)})) + } + require.NoError(t, first.Close()) + + e := openEmbedded(t, dir) + assert.Equal(t, int64(8<<20), e.MaxBytes("acme"), "the budget is read back") + + got := make(chan byte, 2) + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 10}) require.NoError(t, err) - raw, err := dlq.GetLastMsgForSubject(ctx, "dlq.events") + stop, _, err := cons.Consume(func(msg *Message) { + _ = msg.Ack() + got <- msg.Data[0] + }, 4) require.NoError(t, err) - assert.Equal(t, []byte("old"), raw.Data, "parked under the tail it arrived on") + t.Cleanup(stop) + for i := range 2 { + select { + case b := <-got: + assert.Equal(t, byte(i), b) + case <-time.After(5 * time.Second): + t.Fatal("the queued rows of a tenant given no budget this boot were not delivered") + } + } + // And a publish to it opens nothing new: the queue is there at its budget. + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(8<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) } diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 2c5de566..34620f85 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -141,24 +141,30 @@ func WithHeader(key, value string) PublishOpt { } } -// ErrQueueFull is returned by Publisher.Publish when the ingest queue is at -// its byte budget and refuses new events — the backpressure signal the API -// turns into a 503 with Retry-After. +// ErrQueueFull is returned by Publisher.Publish when the topic's tenant's +// ingest queue refuses new events — it is at its byte budget, or the tenant +// has no queue open yet — the backpressure signal the API turns into a 503 +// with Retry-After. var ErrQueueFull = errors.New("ingest queue is full") // Publisher appends events to the ingest queue. type Publisher interface { - // Publish stores data as one event on topic. ErrQueueFull when the queue - // is at its byte budget. + // Publish stores data as one event on topic, in the ingest queue of the + // topic's tenant. ErrQueueFull when that queue is at its byte budget, or + // the tenant has no queue open yet (see Broker.SetMaxBytes). Publish(ctx context.Context, topic Topic, data []byte, opts ...PublishOpt) error Close() error } // Subscriber delivers every event on the ingest queue, across all tenants -// and topics. +// and topics: each tenant's in the order it was published, and different +// tenants' concurrently. type Subscriber interface { // Subscribe registers a handler for incoming events under a durable - // consumer named consumerName. + // consumer named consumerName, held on every tenant's queue — those + // opened after Subscribe included. The handler runs on one delivery + // goroutine per tenant, one message at a time, so it must be safe to + // call concurrently for different tenants. // // CONTRACT: If the handler intends to return an error to trigger automatic // redelivery, it MUST NOT manually call msg.Ack() or msg.Nak() beforehand. @@ -181,27 +187,32 @@ type ConsumerConfig struct { // AckWait is the redelivery timeout: a message not acked within it is // delivered again. AckWait time.Duration - // MaxAckPending caps unacked messages broker-side; delivery pauses when - // hit (backpressure). + // MaxAckPending caps unacked messages broker-side, per tenant: delivery + // of a tenant's events pauses when that tenant's unacked ones hit it + // (backpressure), and no other tenant's does. MaxAckPending int } // Consumer is a live durable consumer created by ConsumerManager. type Consumer interface { - // Consume delivers each message to handler on the client's delivery - // goroutine, so a handler that blocks holds delivery back — that is the - // backpressure the ingest worker relies on. Up to prefetch messages are - // fetched ahead (0 = the client default). The returned stop asks delivery - // to end and returns without waiting: a handler invocation already in - // flight, or one for a message already queued client-side, may still run - // after stop returns, so a handler must not write to anything the caller - // tears down right after stopping. + // Consume delivers each message to handler on a delivery goroutine of + // its tenant's: one per tenant, so a tenant's messages arrive in order, + // one at a time, while different tenants' arrive concurrently — handler + // must be safe for that. A handler that blocks holds back its tenant's + // delivery — that is the backpressure the ingest worker relies on. About + // prefetch messages are fetched ahead across the tenants together, at + // least one per tenant (0 = the client default, per tenant). The returned + // stop asks delivery to end and returns without waiting: a handler + // invocation already in flight, or one for a message already queued + // client-side, may still run after stop returns, so a handler must not + // write to anything the caller tears down right after stopping. // // Delivery can also end on its own after Consume has returned: the broker // or the client gives up on the consumer (it was deleted, the connection - // closed). That is reported on failed — exactly one error, and nothing - // once stop has been called — because no message will ever arrive to say - // so. A caller that ignores failed waits forever on a dead consumer. + // closed), or a tenant's queue opened later could not be joined. That is + // reported on failed — exactly one error, and nothing once stop has been + // called — because no message will ever arrive to say so. A caller that + // ignores failed waits forever on a dead consumer. Consume(handler func(msg *Message), prefetch int) (stop func(), failed <-chan error, err error) } @@ -209,7 +220,8 @@ type Consumer interface { // broker's reason when it gave one. var ErrDeliveryEnded = errors.New("consumer delivery ended") -// ConsumerManager creates durable consumers on the ingest queue. A delivered +// ConsumerManager creates durable consumers on the ingest queue, held on +// every tenant's queue — those opened later included. A delivered // Message.Ctx is the ctx given to CreateConsumer: unlike Subscriber, the // consumer path does not extract the trace context carried in the message // headers, because its one consumer (the ingest worker) batches across @@ -220,51 +232,53 @@ type ConsumerManager interface { // DeadLetterer parks messages on the dead-letter queue. type DeadLetterer interface { - // DeadLetter stores msg's data on the dead-letter queue under msg's topic, - // with the headers the options set. It does not ack msg: the caller acks - // once the parking is confirmed, so a failure here leaves the original to - // be redelivered. + // DeadLetter stores msg's data on the dead-letter queue of msg's tenant, + // under msg's topic, with the headers the options set. It does not ack + // msg: the caller acks once the parking is confirmed, so a failure here + // leaves the original to be redelivered. DeadLetter(ctx context.Context, msg *Message, opts ...PublishOpt) error } -// DeadLetterCounts is what is parked on the dead-letter queue. +// DeadLetterCounts is what is parked on one tenant's dead-letter queue. type DeadLetterCounts struct { - // Tables maps table name → parked messages, for the tables asked about, - // summed across tenants: one queue serves every tenant until each has its - // own (#583 story 5b), so one count covers them all. Scope is - // not broken out yet (it is inert until #235): a message parked under a - // scoped topic counts under "table.scope", not under its table. + // Tables maps table name → parked messages, for the tables asked about. + // Scope is not broken out yet (it is inert until #235): a message parked + // under a scoped topic counts under "table.scope", not under its table. Tables map[string]uint64 - // Total is every parked message, whatever the filter. + // Total is every parked message of the tenant, whatever the filter. Total uint64 } // ErrNoDeadLetterQueue is returned by DeadLetterStats.DeadLetterCounts when -// the dead-letter queue does not exist (nothing can have been parked). Any -// other failure to read it is a plain error. +// the tenant has no dead-letter queue (nothing can have been parked for it). +// Any other failure to read it is a plain error. var ErrNoDeadLetterQueue = errors.New("dead-letter queue not found") -// DeadLetterStats reports on the dead-letter queue. +// DeadLetterStats reports on the dead-letter queues. type DeadLetterStats interface { - // DeadLetterCounts counts parked messages per table; a non-empty table - // narrows Tables to that one (its unscoped messages, under any tenant — - // see DeadLetterCounts.Tables). - DeadLetterCounts(ctx context.Context, table string) (DeadLetterCounts, error) + // DeadLetterCounts counts tenant id's parked messages per table — a + // tenant served, rejected, or removed alike, for as long as its queue is + // kept. A non-empty table narrows Tables to that one (its unscoped + // messages). + DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) } // ErrConsumerNotFound is returned by Purger.PurgeAcked when the named -// consumer does not exist (yet). +// consumer does not exist (yet) on a tenant's queue. var ErrConsumerNotFound = errors.New("consumer not found") // Purger reclaims ingest-queue storage. type Purger interface { - // PurgeAcked removes the ingest events that are BOTH acknowledged by the - // named durable consumer (everything before its first unacked event) AND - // stored before olderThan. Either bound alone keeps the event: unacked - // events are not yet written, and recent ones are still needed for replay. - // Reports whether anything was removed. ErrConsumerNotFound when the - // consumer has not been created. - PurgeAcked(ctx context.Context, consumer string, olderThan time.Time) (purged bool, err error) + // PurgeAcked removes, from each tenant's ingest queue, the events that + // are BOTH acknowledged by the named durable consumer (everything before + // its first unacked event) AND stored before that tenant's cutoff in + // olderThan. Either bound alone keeps the event: unacked events are not + // yet written, and recent ones are still needed for replay. A tenant + // olderThan does not name — one no longer served — keeps no history: + // everything it has acknowledged goes. Reports whether anything was + // removed. ErrConsumerNotFound when the consumer has not been created on + // some tenant's queue; the other tenants' are purged all the same. + PurgeAcked(ctx context.Context, consumer string, olderThan map[tenant.ID]time.Time) (purged bool, err error) } // Replayer re-delivers stored events for SSE gap-fill. @@ -278,7 +292,7 @@ type Replayer interface { } // Broker is everything the process wiring needs from the MQ: every interface -// above plus the lifecycle and the byte budget. EmbeddedNATS is the one +// above plus the lifecycle and the byte budgets. EmbeddedNATS is the one // implementation; internal/app depends on this, not on it. type Broker interface { Publisher @@ -288,15 +302,17 @@ type Broker interface { DeadLetterStats Purger Replayer - // SetMaxBytes applies a new byte budget (the hot-reloadable - // mq.max_bytes_gb) to the queues as a whole — how it is split between - // them is the implementation's. On an error the implementation restores - // the previous budget where it can (best effort: the error says when it - // could not, and a canceled ctx abandons the restore too), and MaxBytes - // keeps reporting the previous budget so the next call retries. - // MaxBytes reports the budget last applied in full. - SetMaxBytes(ctx context.Context, maxBytes int64) error - MaxBytes() int64 + // SetMaxBytes applies tenant id's byte budget (its hot-reloadable + // mq.max_bytes_gb) to that tenant's queues — how it is split between them + // is the implementation's — opening them if the tenant has none yet. No + // other tenant's queues are touched. On an error the implementation + // restores the previous budget where it can (best effort: the error says + // when it could not, and a canceled ctx abandons the restore too), and + // MaxBytes keeps reporting the previous budget so the next call retries. + // MaxBytes reports the budget last applied in full for id, 0 when none + // has been. + SetMaxBytes(ctx context.Context, id tenant.ID, maxBytes int64) error + MaxBytes(id tenant.ID) int64 // Stats reports the broker counters the system gauges observe. Stats() (observability.MQStats, error) } diff --git a/internal/mq/subject.go b/internal/mq/subject.go index 8d67ace4..489ad241 100644 --- a/internal/mq/subject.go +++ b/internal/mq/subject.go @@ -12,25 +12,53 @@ import ( // The embedded broker's naming. Private to this package: everything else // addresses events by Topic. const ( - // ingestStream / dlqStream are the JetStream stream names. Hardcoded — the - // embedded NATS server is private to the WaveHouse process, so there is - // nothing to namespace against. - ingestStream = "WAVEHOUSE" - dlqStream = "WAVEHOUSE_DLQ" + // Each tenant's queue is a pair of JetStream streams named after it: + // INGEST_ and DLQ_. The prefixes differ in their first + // letter, so no tenant id makes one kind's name the other's, and the + // tenant grammar (tenant.Parse: letters, digits, '_' and '-', at most + // tenant.MaxLen bytes) keeps every name inside JetStream's. No namespacing + // beyond that: the embedded server is private to the WaveHouse process. + ingestStreamPrefix = "INGEST_" + dlqStreamPrefix = "DLQ_" + + // legacyIngestStream / legacyDLQStream are the one pair an earlier build + // kept for every tenant. Their subjects (ingest.> and dlq.>) overlap every + // tenant's, and JetStream refuses a stream whose subjects overlap + // another's, so NewEmbedded deletes them. + legacyIngestStream = "WAVEHOUSE" + legacyDLQStream = "WAVEHOUSE_DLQ" // A topic's subject is .
[.]: the tenant id // verbatim — its grammar (tenant.Parse) admits only letters, digits, '_' // and '-', so it is one token as it is — then the table and scope each // as one encoded token. Tenant first so one wildcard selects a tenant's - // traffic (ingest.acme.>). The same topic has the same tail on both - // streams, so parking a message on the DLQ is a prefix swap. + // traffic (ingest.acme.>), which is what the tenant's streams hold. The + // same topic has the same tail on both kinds, so parking a message on the + // dead-letter queue is a prefix swap. ingestPrefix = "ingest." dlqPrefix = "dlq." - - ingestAll = ingestPrefix + ">" // every topic on the ingest stream - dlqAll = dlqPrefix + ">" // every topic on the DLQ stream ) +// ingestStreamName / dlqStreamName name tenant id's two streams. +func ingestStreamName(id tenant.ID) string { return ingestStreamPrefix + string(id) } +func dlqStreamName(id tenant.ID) string { return dlqStreamPrefix + string(id) } + +// tenantSubjects is every subject of tenant id's under prefix: what its +// stream of that kind holds. +func tenantSubjects(prefix string, id tenant.ID) string { return prefix + string(id) + ".>" } + +// streamTenant recovers the tenant a stream name carries under prefix, false +// for any other name: a stream of the other kind, a legacy one, or a name no +// tenant id could have produced. +func streamTenant(prefix, name string) (tenant.ID, bool) { + rest, ok := strings.CutPrefix(name, prefix) + if !ok { + return "", false + } + id, err := tenant.Parse(rest) + return id, err == nil +} + // encodeToken converts any table or scope name into a safe, single NATS // subject token. It preserves alphanumerics and underscores, but // percent-encodes everything else (so '.', ' ', '*' and '>' can never split @@ -66,28 +94,28 @@ func subject(prefix string, t Topic) (string, error) { } // topicKey is the tail of a subject carrying prefix — the key() of the topic -// it was published on, or a one-token tail written before the tenant led the -// subject (see parseTopicKey). A trim, no decoding. +// it was published on. A trim, no decoding. func topicKey(prefix, subj string) string { return strings.TrimPrefix(subj, prefix) } +// keyTenant is the tenant a topic key leads with — the token that decides +// which tenant's stream its subject lands in — whether or not the rest of +// the key parses. +func keyTenant(key string) (tenant.ID, bool) { + first, _, _ := strings.Cut(key, ".") + id, err := tenant.Parse(first) + return id, err == nil +} + // parseTopicKey recovers the Topic from a subject tail. Three tokens are -// tenant, table and scope; two are tenant and table. One token is the form -// this package wrote before the tenant led the subject (#583 story 5) and -// reads as tenant.Default's table: every event of that era was the default -// tenant's, and the durable consumers still deliver them after the upgrade, -// as the dead-letter queue still holds them. A tail this package could not -// have written — more tokens, a token that does not decode, a tenant outside -// the grammar — cannot be split reliably, so the whole of it becomes the -// table of no tenant rather than being dropped. +// tenant, table and scope; two are tenant and table. A tail this package +// could not have written — one token, more than three, a token that does not +// decode, a tenant outside the grammar — cannot be split reliably, so the +// whole of it becomes the table of no tenant rather than being dropped. func parseTopicKey(tail string) Topic { parts := strings.Split(tail, ".") switch len(parts) { - case 1: - if table, err := decodeToken(parts[0]); err == nil && table != "" { - return Topic{Tenant: tenant.Default, Table: table} - } case 2, 3: id, idErr := tenant.Parse(parts[0]) table, tableErr := decodeToken(parts[1]) diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index e4b8bb82..67536e4e 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -1,8 +1,10 @@ package mq import ( + "strings" "testing" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" ) @@ -126,15 +128,6 @@ func TestTopicKey_IsInjective(t *testing.T) { assert.Equal(t, Topic{Tenant: "0", Table: "a", Scope: "b"}.key(), Topic{Tenant: "0", Table: "a", Scope: "b"}.key()) } -// The form written before the tenant led the subject (#583 story 5) is the -// default tenant's: it is what the durable consumers deliver across the -// upgrade, and what the dead-letter queue keeps holding after it. -func TestParseTopicKey_PreTenantTailIsTheDefaultTenants(t *testing.T) { - t.Parallel() - assert.Equal(t, Topic{Tenant: "0", Table: "events"}, parseTopicKey("events")) - assert.Equal(t, Topic{Tenant: "0", Table: "default.clicks"}, parseTopicKey("default%2Eclicks")) -} - func TestParseTopicKey_ForeignTailKeepsItself(t *testing.T) { t.Parallel() // Subjects this package did not write still yield one usable topic, of @@ -144,9 +137,50 @@ func TestParseTopicKey_ForeignTailKeepsItself(t *testing.T) { "0.bad%2Gtoken", // a token that does not decode "a%2Eb.events", // a tenant outside the grammar ".events", // a topic whose tenant was never set + "events", // one token: no tenant leads it "bad%2G", // one token that does not decode } { assert.Equal(t, Topic{Table: tail}, parseTopicKey(tail), tail) } assert.Equal(t, Topic{}, parseTopicKey("")) } + +// The tenant a key leads with picks the stream its subject lands in, so it +// is read off the first token whatever the rest of the key holds. +func TestKeyTenant(t *testing.T) { + t.Parallel() + for key, want := range map[string]tenant.ID{"acme.t": "acme", "a.b.c.d": "a", "0.bad%2G": "0"} { + id, ok := keyTenant(key) + assert.True(t, ok, key) + assert.Equal(t, want, id, key) + } + for _, key := range []string{"", ".events", "a%2Eb.events"} { + _, ok := keyTenant(key) + assert.False(t, ok, key) + } +} + +// Every tenant's two streams have names of their own: no id makes one +// kind's name another stream's, none is a stream an earlier build shared, +// and each name gives its tenant back. +func TestStreamNames_NeverCollide(t *testing.T) { + t.Parallel() + ids := []tenant.ID{"0", "acme", "DLQ", "DLQ_acme", "INGEST", "INGEST_acme", "_", "-", "WAVEHOUSE", tenant.ID(strings.Repeat("a", tenant.MaxLen))} + seen := map[string]tenant.ID{legacyIngestStream: "", legacyDLQStream: ""} + for _, id := range ids { + for prefix, name := range map[string]string{ingestStreamPrefix: ingestStreamName(id), dlqStreamPrefix: dlqStreamName(id)} { + other, dup := seen[name] + assert.False(t, dup, "%s names a stream of %q's too", name, other) + seen[name] = id + back, ok := streamTenant(prefix, name) + assert.True(t, ok, name) + assert.Equal(t, id, back, name) + } + } + for _, name := range []string{legacyIngestStream, legacyDLQStream, "INGEST_a.b", "DLQ_"} { + for _, prefix := range []string{ingestStreamPrefix, dlqStreamPrefix} { + _, ok := streamTenant(prefix, name) + assert.False(t, ok, "%s is no tenant's %s stream", name, prefix) + } + } +} diff --git a/internal/settings/settings.go b/internal/settings/settings.go index d0f836e5..55ec089d 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -163,10 +163,10 @@ type TableDedupe struct { } // DLQConfig gates the Dead Letter Queue: whether a row that still fails -// after the row-by-row isolation retry is parked on the WAVEHOUSE_DLQ stream -// (and its original acked) or left unacked to be redelivered indefinitely. -// The stream itself always exists — it is an empty limits-policy stream -// until something lands on it — so the switch is purely behavioral and +// after the row-by-row isolation retry is parked on the tenant's dead-letter +// queue (and its original acked) or left unacked to be redelivered +// indefinitely. The queue exists from the moment the tenant is first served — +// empty until something lands on it — so the switch is purely behavioral and // resolves per table through the same override cascade as dedupe. type DLQConfig struct { Enabled *bool `json:"enabled"` @@ -219,14 +219,15 @@ type StreamConfig struct { GapWindowMinutes *int `json:"gap_window_minutes"` } -// MQConfig sizes the embedded JetStream streams on disk. +// MQConfig sizes the tenant's message queue on disk. type MQConfig struct { - // MaxBytesGB caps the WAVEHOUSE ingest stream (the DLQ stream gets a - // tenth of it). Must be >= 1. A reload updates the live streams in - // place: growing takes effect immediately; shrinking below what is - // currently buffered makes the ingest stream refuse new publishes - // (DiscardNew → 503 backpressure) until the worker drains it — nothing - // already buffered is dropped. + // MaxBytesGB caps the tenant's ingest queue (its dead-letter queue gets a + // tenth of it). Must be >= 1. A reload updates the live queues in place: + // growing takes effect immediately; shrinking below what is currently + // buffered makes the ingest queue refuse new publishes (DiscardNew → 503 + // backpressure) until the worker drains it — nothing already buffered is + // dropped — and a dead-letter queue holding more than a tenth of the new + // budget keeps what it holds rather than dropping its oldest rows. MaxBytesGB *int `json:"max_bytes_gb"` } diff --git a/internal/settings/store.go b/internal/settings/store.go index d1493917..f68a2fbf 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -193,7 +193,7 @@ func (s *Store) GapWindow() time.Duration { return time.Duration(*s.doc().Config.Stream.GapWindowMinutes) * time.Minute } -// MQMaxBytes returns the ingest stream's disk budget in bytes. +// MQMaxBytes returns the disk budget of the tenant's ingest queue in bytes. func (s *Store) MQMaxBytes() int64 { return int64(*s.doc().Config.MQ.MaxBytesGB) << 30 } diff --git a/internal/stream/subscriber.go b/internal/stream/subscriber.go index 54883492..5952aa1a 100644 --- a/internal/stream/subscriber.go +++ b/internal/stream/subscriber.go @@ -55,15 +55,17 @@ type Subscriber struct { // which reads nothing here but may run alongside the fan-out. // // It does NOT make Hub.deliver's check→send→record sequence atomic, and - // deliver does not need it to be: Broadcast runs on ONE goroutine — the - // single jetstream Consume callback the hub bridge registers in - // internal/app, invoked inline per message — so no two events race - // to announce the same connection's columns. A future change that fans - // Broadcast out across goroutines must hold a lock across that whole - // sequence, or two events will both send an announcement (harmless) while a - // third slips a row between a check and its record (not harmless: the client - // zips it against the previous list). Replay does not touch this field at - // all — it tracks drift in its own closure; see Hub.ReplayProjector. + // deliver does not need it to be: the hub bridge registered in + // internal/app calls Broadcast inline per message, on one delivery + // goroutine per tenant (mq.Subscriber), and a connection subscribes to one + // tenant's topic — so all of its events come from that one goroutine, and + // no two events race to announce the same connection's columns. A future + // change that fans one tenant's Broadcasts out across goroutines must hold + // a lock across that whole sequence, or two events will both send an + // announcement (harmless) while a third slips a row between a check and its + // record (not harmless: the client zips it against the previous list). + // Replay does not touch this field at all — it tracks drift in its own + // closure; see Hub.ReplayProjector. schemaMu sync.Mutex // lastSchema is the signature of the column list most recently announced to // this connection ("" ⇒ none yet). Rows travel positionally, so a client that diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index 1b8cf358..44f3ebe6 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -221,10 +221,10 @@ type MockPurger struct { // PurgeCall records one PurgeAcked call. type PurgeCall struct { Consumer string - OlderThan time.Time + OlderThan map[tenant.ID]time.Time } -func (m *MockPurger) PurgeAcked(_ context.Context, consumer string, olderThan time.Time) (bool, error) { +func (m *MockPurger) PurgeAcked(_ context.Context, consumer string, olderThan map[tenant.ID]time.Time) (bool, error) { m.mu.Lock() defer m.mu.Unlock() m.Calls = append(m.Calls, PurgeCall{Consumer: consumer, OlderThan: olderThan}) @@ -233,13 +233,21 @@ func (m *MockPurger) PurgeAcked(_ context.Context, consumer string, olderThan ti // ── Mock mq.DeadLetterStats ────────────────────────────────────── -// MockDeadLetterStats implements mq.DeadLetterStats with a canned answer. +// MockDeadLetterStats implements mq.DeadLetterStats with a canned answer, +// recording the tenant and table each call asked about. type MockDeadLetterStats struct { Counts mq.DeadLetterCounts Err error + + mu sync.Mutex + Tenant tenant.ID + Table string } -func (m *MockDeadLetterStats) DeadLetterCounts(context.Context, string) (mq.DeadLetterCounts, error) { +func (m *MockDeadLetterStats) DeadLetterCounts(_ context.Context, id tenant.ID, table string) (mq.DeadLetterCounts, error) { + m.mu.Lock() + defer m.mu.Unlock() + m.Tenant, m.Table = id, table return m.Counts, m.Err } diff --git a/internal/testutil/testutil.go b/internal/testutil/testutil.go index 371cdcd4..db119686 100644 --- a/internal/testutil/testutil.go +++ b/internal/testutil/testutil.go @@ -14,6 +14,7 @@ import ( "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/discovery" + "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -39,6 +40,24 @@ func NewTestSchemaRegistry(t testing.TB, tables []*discovery.TableSchema) *disco // hardcoding the same literal twice. const TestServerVersion = "24.8.1.1" +// NewEmbeddedMQ starts the embedded broker over a temporary directory, closed +// by the test framework, with a queue open for each of tenants — +// tenant.Default when none is named — at maxBytes: a tenant has a queue once +// its budget is applied, as the wiring does for every tenant it serves. +func NewEmbeddedMQ(t testing.TB, maxBytes int64, tenants ...tenant.ID) *mq.EmbeddedNATS { + t.Helper() + emb, err := mq.NewEmbedded(t.TempDir()) + require.NoError(t, err) + t.Cleanup(func() { _ = emb.Close() }) + if len(tenants) == 0 { + tenants = []tenant.ID{tenant.Default} + } + for _, id := range tenants { + require.NoError(t, emb.SetMaxBytes(context.Background(), id, maxBytes)) + } + return emb +} + // schemaConn is a mock driver.Conn serving exactly the queries Refresh issues: // the SELECT timezone() (always "UTC") and SELECT version() probes, the // system.columns scan (rows synthesized from tables), and the system.tables DDL From 8ad98656e412709b5ab879379de50cd49321b493 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 18:25:18 -0400 Subject: [PATCH 002/108] fix(mq): reopen detached from the request, and the upgrade runbook swept --- docs/src/content/docs/api.md | 6 +-- docs/src/content/docs/architecture.md | 9 +++-- docs/src/content/docs/deployment.md | 8 ++-- docs/src/content/docs/durability.md | 4 +- docs/src/content/docs/ingest-pipeline.md | 10 ++--- docs/src/content/docs/sdk/streaming.md | 2 +- docs/src/content/docs/settings-directory.mdx | 4 +- internal/ingest/worker.go | 25 ++++++------ internal/ingest/worker_test.go | 6 +-- internal/mq/embedded.go | 13 +++++- internal/mq/embedded_test.go | 42 ++++++++++++++++++++ internal/settings/settings.go | 7 ++-- 12 files changed, 96 insertions(+), 40 deletions(-) diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 1634aaab..ad712b82 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -274,7 +274,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error | -| 503 | `{"error":"service unavailable"}` | NATS JetStream stream full (backpressure). Response includes `Retry-After: 30` header. | +| 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** @@ -385,7 +385,7 @@ A `200` is returned whenever the body was read and the records were processed | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch | -| 503 | `{"error":"service unavailable"}` | NATS JetStream full (backpressure) mid-batch; includes `Retry-After: 30` | +| 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30` | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] @@ -876,7 +876,7 @@ Three values, where the envelope above has four: this is the frame a role restri ## Dead Letter Queue (DLQ) -When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the tenant's own DLQ NATS stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format` (a pre-v2 message has no `format` field at all, which is how it presents here), or `columns` and `row` that do not pair — is parked without ever reaching a table batch, which is what an operator sees after upgrading across the wire change without draining first. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. +When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the tenant's own DLQ NATS stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format`, or `columns` and `row` that do not pair — is parked without ever reaching a table batch. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. Use `GET /v1/ops/dlq/stats` to monitor DLQ depth, per tenant (`?tenant=`). diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54..92508dac 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -190,7 +190,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `tenant/` — Tenant Identifier -- **tenant.go** — `ID`, a validated string (never a number: a 19-digit id already rounds as a float64), and `Parse`, the one grammar that makes an id safe both as a folder name and as a message-queue subject token: ASCII letters, digits, `_`, `-`, at most `MaxLen` (64) bytes. `Default` (`"0"`) is the tenant a request without the header resolves to; `Header` is `X-Tenant-ID`. The package imports nothing from the rest of the repository, so any package can name a tenant. HTTP handlers receive the tenant as its resolved `*settings.Store`, which knows its id (`Store.Tenant`) for the topics they publish and subscribe on; the stream hub and the ingest worker read each event's tenant off its `mq.Topic` — the leading subject token — and their settings getters take it as a parameter, which `internal/app` resolves through the registry; the sweeper folds over the tenants served; each served tenant has a schema registry of its own, built with its id (story 6). +- **tenant.go** — `ID`, a validated string (never a number: a 19-digit id already rounds as a float64), and `Parse`, the one grammar that makes an id safe both as a folder name and as a message-queue subject token: ASCII letters, digits, `_`, `-`, at most `MaxLen` (64) bytes. `Default` (`"0"`) is the tenant a request without the header resolves to; `Header` is `X-Tenant-ID`. The package imports nothing from the rest of the repository, so any package can name a tenant. HTTP handlers receive the tenant as its resolved `*settings.Store`, which knows its id (`Store.Tenant`) for the topics they publish and subscribe on; the stream hub and the ingest worker read each event's tenant off its `mq.Topic` — the leading subject token — and their settings getters take it as a parameter, which `internal/app` resolves through the registry; the sweeper hands the MQ each served tenant's own gap window (`gapWindows`); each served tenant has a schema registry of its own, built with its id (story 6). ### `chconn/` — ClickHouse Connection Pools @@ -228,7 +228,7 @@ Client POST /v1/ingest?table={table} field is published un-deduped + logged/counted, or rejected under require_id) → Publish to NATS JetStream (ingest.{tenant}.{table}) → 200 OK returned immediately - → (If NATS stream is full: 503 + Retry-After header) + → (If the tenant's NATS stream is full, or not open: 503 + Retry-After header) Ingest worker pipeline (StartIngestWorker): ← JetStream pull consumer (buffer-consumer) on ingest.> @@ -251,9 +251,10 @@ Ingest worker pipeline (StartIngestWorker): a no/invalid-token request (resolved to default_role, not admin in a production config) cannot reach the proxy.) -Active Sweeper (async goroutine, every 60s): +Active Sweeper (async goroutine, every 60s), on each tenant's stream: → Read buffer consumer's AckFloor (highest contiguous ACKed seq) - → Binary search for first message within the gap window (the longest among the tenants served) + → Binary search for first message within that tenant's own gap window + (none for a tenant no longer served) → Purge target = MIN(ack_floor + 1, gap_window_seq) → Purge all messages below target from JetStream ``` diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 095b4990..b2bac204 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -421,9 +421,9 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres The NATS envelope changed shape in this release: the row now travels positionally, with `format`, `columns` and `row` replacing `data`. **The new worker cannot read a message published by an older version** — it carries no `format`, so there is no way to say which value belongs to which column. -This affects the streaming surface too, and more quietly. SSE gap-fill (`?since=` / `Last-Event-ID`) reads the same stream, and the hub refuses a pre-v2 envelope on the same missing `format` the worker does — it is withheld from every role with **no error and no frame**, and the stream side files no DLQ entry of its own — the worker's copy of the same message is what lands in `dlq.{table}` (next paragraph). Because the stream keeps ACKed messages until the sweeper purges past `stream.gap_window_minutes` (15 by default), this outlives a *correct* drain: for that window, any replay spanning the upgrade silently omits the pre-upgrade events. Clients that need them should backfill over REST. +This affects the streaming surface too, and more quietly. SSE gap-fill (`?since=` / `Last-Event-ID`) replays from the queue, and the upgrade deletes the old one (next paragraph) — the acknowledged history kept for replay included — so this outlives a *correct* drain: any replay spanning the upgrade silently omits the pre-upgrade events, with **no error and no frame**. Clients that need them should backfill over REST. -On the worker side the outcome depends on the DLQ. **With the DLQ enabled for the table**, the message is parked on `dlq.{table}` with `X-DLQ-*` headers and is recoverable by hand — but re-ingest each parked envelope's inner `data` object as a fresh `POST /v1/ingest`; republishing the envelope as-is onto `ingest.{tenant}.{table}` fails the same `format` check and simply re-parks it. **With the DLQ switched off for the table, it is permanently lost**: acked and dropped with an `ERROR` log and a `wavehouse_ingest_poison_total` increment carrying `disposition="dropped"`, unrecoverable from either the ingest stream or the DLQ, because a message that can never insert must not redeliver forever. Draining first is cheaper than a manual replay, and it is the only option at all where the DLQ is off. +**The upgrade does not carry the old queue over at all.** Boot deletes the earlier build's queue and dead-letter queue (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) and everything in them, logging a `WARN` with each one's message count: an event the old build had not yet inserted, and a row it had already parked, do not survive the upgrade. Draining first is the only way to keep them, and anything parked has to be replayed before the upgrade — re-ingest each parked envelope's inner `data` object as a fresh `POST /v1/ingest`. Three audits belong **before** the drain, because none of them announces itself afterwards: @@ -435,10 +435,10 @@ To drain before upgrading: 1. **Stop the producers**, or cut `/v1/ingest` at the reverse proxy. Nothing new should enter the stream. 2. **Wait for the in-flight batches to flush.** A table's batch closes on size or after `maxWait` (5s by default), so a few seconds after the last write is enough; give it longer if ClickHouse is slow or retrying. -3. **Confirm nothing is left unconsumed** before swapping binaries. Not that the stream is empty: it is dual-use, and deliberately retains ACKed messages for the SSE replay window, so a non-zero depth right after a clean drain is expected. Rows landing in ClickHouse is a success signal, **not proof the queue is drained** — when the DLQ is off for a table, a row that fails its retry is skipped without being acked, so NATS keeps redelivering it while its neighbors land. Check that nothing is still failing or redelivering, and note which signal covers which case: [`GET /v1/ops/dlq/stats`](/api#get-v1opsdlqstats--dlq-statistics) is non-zero only where the DLQ is **on**; `wavehouse_ingest_poison_total` counts unreadable envelopes on **either** setting, told apart by its `disposition` label (`parked` / `dropped`); and for a twice-failed row with the DLQ off — the case just described — the **only** signal is the `ERROR` log (`isolated bad row, DLQ disabled for table`). A clean `dlq/stats` with the DLQ off proves nothing. There is no queue-depth gauge today ([#544](https://github.com/Wave-RF/WaveHouse/issues/544) tracks the related in-flight accounting), and `wavehouse_nats_in_msgs_total` going flat is a supporting signal rather than a guarantee. Enabling the DLQ is not itself a drain — replay from `dlq.{table}` is manual. +3. **Confirm nothing is left unconsumed** before swapping binaries. Not that the stream is empty: it is dual-use, and deliberately retains ACKed messages for the SSE replay window, so a non-zero depth right after a clean drain is expected. Rows landing in ClickHouse is a success signal, **not proof the queue is drained** — when the DLQ is off for a table, a row that fails its retry is skipped without being acked, so NATS keeps redelivering it while its neighbors land. Check that nothing is still failing or redelivering, and note which signal covers which case: [`GET /v1/ops/dlq/stats`](/api#get-v1opsdlqstats--dlq-statistics) is non-zero only where the DLQ is **on**; `wavehouse_ingest_poison_total` counts unreadable envelopes on **either** setting, told apart by its `disposition` label (`parked` / `dropped`); and for a twice-failed row with the DLQ off — the case just described — the **only** signal is the `ERROR` log (`isolated bad row, DLQ disabled for table`). A clean `dlq/stats` with the DLQ off proves nothing. There is no queue-depth gauge today ([#544](https://github.com/Wave-RF/WaveHouse/issues/544) tracks the related in-flight accounting), and `wavehouse_nats_in_msgs_total` going flat is a supporting signal rather than a guarantee. Enabling the DLQ is not itself a drain — replay from `dlq.{table}` is manual, and has to happen before the upgrade deletes it. 4. **Upgrade**, then re-enable ingest. -If you skipped the drain, check `wavehouse_ingest_poison_total`, which counts both — `disposition="parked"` is recoverable from `dlq.{table}`, `disposition="dropped"` is gone — see [Dead Letter Queue](#dead-letter-queue-dlq) below. +If you skipped the drain, the boot's `WARN` line for each deleted stream (`deleted the stream an earlier build kept for every tenant together`) says how many messages went with it. ## Dead Letter Queue (DLQ) diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 98a8c9c2..353608f2 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -81,7 +81,7 @@ Read the measured p99 against these bands, which track WaveHouse's `SyncAlways` | 1–5 ms | **Good** | | 5–50 ms | **Workable** — watch bursty load | | 50 ms – 1 s | **Marginal** — relax durability once `mq.sync_interval` ([#139](https://github.com/Wave-RF/WaveHouse/issues/139)) lands, or move to faster storage | -| > 1 s | **Broken** — `create stream` will time out under load; fix the storage substrate | +| > 1 s | **Broken** — opening a tenant's queue (`open dlq stream`) will time out under load; fix the storage substrate | :::caution[macOS `fsync` lies by default] A plain `fsync()` on macOS returns once data is in the drive's volatile cache — it does **not** force a flush to NAND; only `fcntl(fd, F_FULLFSYNC)` does (NATS, Postgres, and SQLite all use it). On a Mac, any per-flush number under ~1 ms is almost certainly not a real flush — the gap between plain `fsync()` and `F_FULLFSYNC` can be ~180× on the same consumer NVMe. `fio` on macOS calls plain `fsync()`, so don't trust Mac `fio` numbers for tail-latency planning. This mostly matters when benchmarking a dev machine; production WaveHouse runs on Linux, where `fio` is honest. @@ -93,7 +93,7 @@ A self-contained `wavehouse storage-check` preflight subcommand that bakes this If you see any of these, benchmark the `/nats` volume as above: -- `create stream: ... context deadline exceeded` at startup. +- `open dlq stream: ... context deadline exceeded` when a tenant's queue first opens, at the first boot or at the reload that adopts the tenant. - Ingest p99 latency in the seconds, or occasional `200`s that take multiple seconds to return. - Intermittent `503 Service Unavailable` from `/v1/ingest` when ClickHouse is healthy (the worker can't drain fast enough because acking is `fsync`-bound). - Flaky CI or load tests that pass on fast storage and fail on a shared/virtualized host. diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index e602e0cf..b504fef4 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -61,7 +61,7 @@ Inserts also pin `input_format_null_as_default=1`. A positional row has one valu ::: :::note[ClickHouse timestamp parsing] -Inserts pin `date_time_input_format=best_effort` — the server default since ClickHouse 26.5, but on older servers the `basic` default rejects the canonical RFC 3339 form's `Z` suffix ([#372](https://github.com/Wave-RF/WaveHouse/issues/372)). The ordinary spellings (zone-less date-times, 9–10-digit Unix-seconds strings) parse identically under both settings. (This is moot for anything still buffered from an older build: a message published before the v2 envelope cannot be read at all — see [Upgrading across the v2 ingest envelope](/deployment#upgrading-across-the-v2-ingest-envelope).) Bare digit-strings of other lengths are the exception: `best_effort` reads them as ClickHouse's calendar/epoch shapes, where `basic` read a plain `DateTime` column's digit string of five or more digits as Unix seconds (shorter runs it rejected outright, where `best_effort` reads `"2026"` as a year): under `best_effort` `"20260711"` stores 2026-07-11, where `basic` stored 1970-08-23. `DateTime64` columns diverge the same way on calendar-shaped runs, and additionally whenever an epoch run's unit doesn't match the column scale (under `basic`, runs longer than 10 digits are ticks at the column's own scale; `best_effort` unit-detects 13/16/19-digit runs as ms/µs/ns). A producer relying on the old `basic` reading changes meaning as soon as this WaveHouse version is deployed — the pin, not a ClickHouse upgrade, is what flips the parse. +Inserts pin `date_time_input_format=best_effort` — the server default since ClickHouse 26.5, but on older servers the `basic` default rejects the canonical RFC 3339 form's `Z` suffix ([#372](https://github.com/Wave-RF/WaveHouse/issues/372)). The ordinary spellings (zone-less date-times, 9–10-digit Unix-seconds strings) parse identically under both settings. (This is moot for anything an older build buffered: the upgrade deletes it — see [Upgrading across the v2 ingest envelope](/deployment#upgrading-across-the-v2-ingest-envelope).) Bare digit-strings of other lengths are the exception: `best_effort` reads them as ClickHouse's calendar/epoch shapes, where `basic` read a plain `DateTime` column's digit string of five or more digits as Unix seconds (shorter runs it rejected outright, where `best_effort` reads `"2026"` as a year): under `best_effort` `"20260711"` stores 2026-07-11, where `basic` stored 1970-08-23. `DateTime64` columns diverge the same way on calendar-shaped runs, and additionally whenever an epoch run's unit doesn't match the column scale (under `basic`, runs longer than 10 digits are ticks at the column's own scale; `best_effort` unit-detects 13/16/19-digit runs as ms/µs/ns). A producer relying on the old `basic` reading changes meaning as soon as this WaveHouse version is deployed — the pin, not a ClickHouse upgrade, is what flips the parse. ::: ## The journey of one event @@ -70,7 +70,7 @@ Inserts pin `date_time_input_format=best_effort` — the server default since Cl sequenceDiagram participant P as POST /v1/ingest participant JS as JetStream - participant CB as Consume callback + participant CB as Consume callback (the tenant's) participant D as dispatchLoop participant TL as tableLoop participant CH as ClickHouse @@ -89,11 +89,11 @@ sequenceDiagram ## Goroutine topology -The design rule is **single-owner state, lock-free**: each piece of mutable state is touched by exactly one goroutine. There are no mutexes in the hot path. +The design rule is **single-owner state, lock-free**: each piece of mutable state is touched by exactly one goroutine. There are no mutexes in the hot path. The one fan-in is at the top: each tenant's stream is delivered on a nats.go goroutine of its own, and they all send into the one `msgChan`, which is safe from all of them at once; everything from `dispatchLoop` down stays single-owner, and a full `msgChan` pauses every tenant's delivery (layer 2 below). ```mermaid flowchart TD - CB["Consume callback
(nats.go goroutine)"] -->|"msgChan (cap maxBatch*2)"| D + CB["Consume callbacks
(one nats.go goroutine per tenant stream)"] -->|"msgChan (cap maxBatch*2)"| D D["dispatchLoop
1 goroutine — owns the routing map
the ONLY ctx watcher — tracked by wg"] D -->|"per-tenant-table chan (cap maxBatch)"| T1["tableLoop: clicks
owns its batch + timer
tracked by tableWg"] D --> T2["tableLoop: events
tracked by tableWg"] @@ -213,7 +213,7 @@ Several layers throttle the pipeline, inner to outer: 1. **`batch`** flushes at `maxBatch` rows or `maxWait`. 2. **`msgChan`** (cap `maxBatch*2`) — when full, the consume callback blocks and delivery pauses. 3. **`pullMaxMessages`** — nats.go's client-side prefetch buffer in front of `msgChan`, shared by the tenants' streams (at least one message each). -4. **`maxAckPending`** — the server suspends a tenant's delivery once this many of its messages are delivered-but-unacked; no other tenant's delivery waits on it. The outermost in-memory bound. +4. **`maxAckPending`** — the server suspends a tenant's delivery once this many of its messages are delivered-but-unacked; no other tenant's delivery waits on it. The outermost in-memory bound, and a per-tenant one: while ClickHouse stalls, the worker can hold up to `maxAckPending` rows for every tenant served. 5. **`MaxBytes` + `DiscardNew`** on each tenant's stream (its `mq.max_bytes_gb` in the [settings directory](/settings-directory#message-queue), resized in place on reload) — when it fills (e.g. ClickHouse is down so nothing acks/purges), that tenant's new publishes are rejected and the API returns 503. | Knob | Default | Meaning / invariant | diff --git a/docs/src/content/docs/sdk/streaming.md b/docs/src/content/docs/sdk/streaming.md index aab6dc6e..94a7e334 100644 --- a/docs/src/content/docs/sdk/streaming.md +++ b/docs/src/content/docs/sdk/streaming.md @@ -116,7 +116,7 @@ A dropped stream reconnects on a jittered exponential backoff, capped at 30s, an :::caution[Resumption is at-least-once, and time-bounded] Delivery across a reconnect is **at-least-once**. The `Last-Event-ID` the client sends is the last event's `received_timestamp`, and the server replays from that instant *inclusively* — so the last event you already saw, and anything sharing its timestamp, arrives again. The SDK does not deduplicate live frames — `liveQuery()` makes one pass at the backfill seam, and only under an ascending order ([#449](https://github.com/Wave-RF/WaveHouse/issues/449)) — so key on `timestamp` plus your own row identity if duplicates matter. -Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. The same silence applies for one gap window after a server upgrade across the v2 ingest envelope: the hub refuses an envelope whose `format` it does not recognize and a pre-v2 message carries none, so a replay spanning that boundary omits them without an error — backfill over REST if you need them. +Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. The same silence applies across a server upgrade to this release: the server deletes the previous release's queue at boot, so a replay spanning the upgrade omits the events published before it, without an error — backfill over REST if you need them. **A column-set change across a gap-fill is a known limitation.** If the table's columns change while you are connected *and* your client replays across that change, live rows arriving after the replay may not be preceded by a fresh `event: schema` frame until the columns next change or you reconnect. The SDK drops a row whose **length** disagrees with the list it was last told, rather than zipping it under the wrong names — so an added or removed column costs you rows, not wrong ones. A **same-length** change is the residual case the arity check cannot see: a `RENAME COLUMN`, or a drop paired with an add, zips values under the wrong names until the next announcement. Reconnecting resynchronizes either way. Full schema-change handling is deferred to the schema-versioning work ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)). ::: diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e7a19b..41e04187 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -213,7 +213,7 @@ The `auth` block is the verifier wiring, minus the secrets. `jwks_url` (absolute A failed batch insert is retried row by row; a row that fails again on its own is a poison row. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the retry, which no row of it could pass, and every row of it is a poison row. `dlq.enabled` (seed default `true`) decides what happens to it, resolved per table (`dlq.tables.
.enabled` → global) at the moment of the failure, so a reload applies to the next poison row: - `true` — the row is published to the tenant's dead-letter stream (`DLQ_{tenant}`) under `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) with the failure in its headers, and its original is acked. Inspect it with `GET /v1/ops/dlq/stats` (admin-only; `?tenant=` names the tenant). -- `false` — the row is left unacked, so NATS redelivers it and it retries until it inserts or the switch is flipped back. For every row the worker **can read**, nothing is ever dropped either way — the choice is *park it* versus *keep retrying*. **One exception, new in this release:** an envelope the worker cannot read *at all* — malformed JSON, an unknown `format` (what a pre-v2 in-flight message looks like), or `columns` and `row` that do not pair — can never insert, so redelivering it forever would wedge the consumer. With the DLQ off for the table it is acked and **dropped**, logged at `ERROR` and counted by `wavehouse_ingest_poison_total` with `disposition="dropped"` (also labeled by `table` and `reason`; an envelope parked on the DLQ carries `disposition="parked"`). See [Ingest Pipeline](/ingest-pipeline) — and drain the ingest queue before upgrading. +- `false` — the row is left unacked, so NATS redelivers it and it retries until it inserts or the switch is flipped back. For every row the worker **can read**, nothing is ever dropped either way — the choice is *park it* versus *keep retrying*. **One exception, new in this release:** an envelope the worker cannot read *at all* — malformed JSON, an unknown `format`, or `columns` and `row` that do not pair — can never insert, so redelivering it forever would wedge the consumer. With the DLQ off for the table it is acked and **dropped**, logged at `ERROR` and counted by `wavehouse_ingest_poison_total` with `disposition="dropped"` (also labeled by `table` and `reason`; an envelope parked on the DLQ carries `disposition="parked"`). See [Ingest Pipeline](/ingest-pipeline) — and drain the ingest queue before upgrading. For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in its ingest queue, redelivered for as long as the tenant is away and stopping the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging that queue. @@ -221,7 +221,7 @@ A tenant's dead-letter stream exists from the moment the tenant is first served ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the worker drains it back under the limit — nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to — so a queue fills until its budget or the disk runs out, whichever comes first; size them together from [Durability & Storage](/durability). +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume. A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. ## Streaming diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index 9af5b1ef..b4defa36 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -117,7 +117,10 @@ const ( // messages are redelivered mid-processing → duplicate inserts). const ( // Server-side cap on a tenant's unacked messages; suspends that tenant's - // delivery when hit (backpressure), and no other tenant's. + // delivery when hit (backpressure), and no other tenant's. The worker holds + // every delivered row until its batch is acked, so while ClickHouse stalls + // it can hold up to maxAckPending rows per tenant: the in-memory bound + // grows with the tenants served. maxAckPending = 10_000 // TODO: raise if NATS delivery becomes the bottleneck // Client prefetch buffer in front of msgChan (was the implicit jetstream @@ -471,13 +474,12 @@ func firstDuplicate(cols []string) (string, bool) { } // parseMsg unmarshals one envelope into a parsedMsg. An envelope the worker can -// never insert is poison — malformed JSON, a row format it doesn't know (which -// is what a pre-v2 envelope looks like: it carries no `format` at all), or +// never insert is poison — malformed JSON, a row format it doesn't know, or // columns and a row it can't pair. Poison is parked on the DLQ rather than -// dropped, so an operator who skipped the documented pre-deploy drain finds -// those rows waiting instead of gone; when the DLQ is off for the table it is -// acked-and-dropped with a counted error, because a message that can never -// insert must not redeliver forever. ok is false either way so the caller skips it. +// dropped, so an operator finds those rows waiting instead of gone; when the +// DLQ is off for the table it is acked-and-dropped with a counted error, +// because a message that can never insert must not redeliver forever. ok is +// false either way so the caller skips it. func (w *IngestWorker) parseMsg(ctx context.Context, m *mq.Message) (parsedMsg, bool) { var envelope EventMessage @@ -493,7 +495,7 @@ func (w *IngestWorker) parseMsg(ctx context.Context, m *mq.Message) (parsedMsg, slog.ErrorContext(ctx, "event envelope declares an unknown row format", "format", envelope.Format, "tenant", id, "table", envelope.TableName) w.rejectPoison(ctx, m, id, envelope.TableName, "unknown_format", - fmt.Sprintf("unknown row format %q (a pre-v2 envelope carries none); drain the ingest queue before upgrading", envelope.Format)) + fmt.Sprintf("unknown row format %q", envelope.Format)) return parsedMsg{}, false } if len(envelope.Columns) == 0 || len(envelope.Row) == 0 { @@ -794,10 +796,9 @@ func (w *IngestWorker) rejectPoison(ctx context.Context, m *mq.Message, id tenan if w.dlqEnabled == nil || w.dlqEnabled(id, tableName) { // Backgrounded on ackWg for the same reason handleSuccess backgrounds its // acks: parkOnDLQ does a DLQ publish AND an fsync-bound DoubleAck, - // and parseMsg runs on the dispatchLoop goroutine. The scenario this whole - // change targets is an operator who skipped the drain, where EVERY backlog - // message is poison — done inline that is one publish plus one fsync per - // message in series, with intake stalled behind it. dispatchLoop adds and + // and parseMsg runs on the dispatchLoop goroutine. When a whole backlog + // is poison, done inline that is one publish plus one fsync per message + // in series, with intake stalled behind it. dispatchLoop adds and // waits on the same goroutine, so each Add still happens-before the Wait. w.ackWg.Go(func() { if w.parkOnDLQ(ctx, m, tableName, detail) { diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index 7fe9130e..935eae77 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -1265,10 +1265,10 @@ func v1Envelope(t *testing.T, table string, data map[string]any) []byte { } // TestParseMsg_PoisonEnvelope_ParkedOnDLQ: an envelope the worker can never -// insert — a pre-v2 message left in the queue across an upgrade, malformed +// insert — one of an unknown format (the pre-v2 shape carries none), malformed // JSON, or columns and a row that can't be paired — is preserved on the DLQ -// rather than dropped, so a missed pre-deploy drain costs an operator a replay -// rather than the rows themselves. +// rather than dropped, so it costs an operator a replay rather than the rows +// themselves. func TestParseMsg_PoisonEnvelope_ParkedOnDLQ(t *testing.T) { t.Parallel() tests := []struct { diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index dbe35fa5..e082915c 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -336,8 +336,13 @@ func (e *EmbeddedNATS) apply(ctx context.Context, id tenant.ID, q *tenantQueue, return fmt.Errorf("open ingest stream: %w", err) } q.ingest, q.maxBytes = true, maxBytes + // The joins run on a budget of their own: a queue that opened but no + // consumer holds fails every consumer (fail), so a slow open must not + // leave them no time. + joinCtx, cancelJoin := context.WithTimeout(ctx, resizeTimeout) + defer cancelJoin() for _, f := range e.consumers { - if err := f.join(resizeCtx, id); err != nil { + if err := f.join(joinCtx, id); err != nil { f.fail(fmt.Errorf("tenant %s: %w: join its queue: %w", id, ErrDeliveryEnded, err)) } } @@ -394,7 +399,13 @@ func (e *EmbeddedNATS) applyDLQ(ctx context.Context, id tenant.ID, q *tenantQueu // publish or park that found one of its streams missing. errNoQueue when no // budget has been asked for the tenant yet: a reload can make a tenant // resolvable an instant before its budget arrives. +// +// It runs detached from ctx's cancellation, bounded by its own timeouts: +// ctx is one caller's — an ingest request — while the queue is every +// consumer's, and a client that goes away between the open and the joins +// would leave a queue no consumer holds, which fails the ingest worker. func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { + ctx = context.WithoutCancel(ctx) e.mu.Lock() defer e.mu.Unlock() q := e.queues[id] diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index fe4fd45e..06120cc5 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -713,6 +713,48 @@ func TestEmbeddedNATS_Publish_OpensTheQueueAtTheLastBudget(t *testing.T) { require.ErrorIs(t, err, jetstream.ErrStreamNotFound) } +// The context a publish reopens a queue under is one client's request, but +// the queue is every consumer's: a client gone before the consumers join must +// not leave a queue that no consumer holds, which the ingest worker would +// report as its delivery ending. So the reopen — joins included — outlives +// the caller's cancellation. +func TestEmbeddedNATS_ReopenOutlivesTheCallersCancellation(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 10}) + require.NoError(t, err) + for _, name := range []string{"INGEST_acme", "DLQ_acme"} { + require.NoError(t, e.js.DeleteStream(ctx, name)) + } + + gone, stop := context.WithCancel(ctx) + stop() + require.NoError(t, e.reopen(gone, "acme")) + + _, err = e.js.Consumer(ctx, "INGEST_acme", "buffer") + require.NoError(t, err, "the consumer joined the reopened queue") + select { + case err := <-cons.(*workerConsumer).failed: + t.Fatalf("the reopen was reported as the consumer's failure: %v", err) + default: + } + got := make(chan byte, 1) + stopConsume, _, err := cons.Consume(func(msg *Message) { + _ = msg.Ack() + got <- msg.Data[0] + }, 4) + require.NoError(t, err) + t.Cleanup(stopConsume) + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte{7})) + select { + case b := <-got: + assert.Equal(t, byte(7), b) + case <-time.After(5 * time.Second): + t.Fatal("the reopened queue is not delivered") + } +} + func TestEmbeddedNATS_PurgeAcked(t *testing.T) { e := newTestEmbedded(t) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 55ec089d..2ce9b120 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -225,9 +225,10 @@ type MQConfig struct { // tenth of it). Must be >= 1. A reload updates the live queues in place: // growing takes effect immediately; shrinking below what is currently // buffered makes the ingest queue refuse new publishes (DiscardNew → 503 - // backpressure) until the worker drains it — nothing already buffered is - // dropped — and a dead-letter queue holding more than a tenth of the new - // budget keeps what it holds rather than dropping its oldest rows. + // backpressure) until the sweeper purges it back under the limit — + // nothing already buffered is dropped — and a dead-letter queue holding + // more than a tenth of the new budget keeps what it holds rather than + // dropping its oldest rows. MaxBytesGB *int `json:"max_bytes_gb"` } From 650a28e20a4f3ab9879e021341cbfa4c2b9b7bd3 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 18:57:43 -0400 Subject: [PATCH 003/108] fix(app): boot opens queues under New's context; docs review fixes --- AGENTS.md | 2 +- docs/src/content/docs/api.md | 6 +-- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 10 ++--- docs/src/content/docs/ingest-pipeline.md | 2 +- docs/src/content/docs/sdk/streaming.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app.go | 5 ++- internal/app/app_test.go | 44 ++++++++++++++++++++ internal/app/wire.go | 26 +++++++----- internal/mq/embedded.go | 7 +++- 11 files changed, 81 insertions(+), 27 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 16595721..dbf19e84 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -56,7 +56,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 3. **Schema-driven ingest** — `POST /v1/ingest?table={table}` takes flat JSON, validated against the discovered schema (unknown fields rejected, types/nullability enforced). No envelope. The **declared `Content-Type` chooses the format and the bytes never do** (arity within the JSON family is still the body's): no declaration, one whose **media type** is unsupported or unparseable, a comma-bearing value that, as a whole, does not parse as one media type, or repeated lines that **disagree**, is a `415` decided *before* the body is read. A malformed *parameter* on a comma-free line never costs the request (`; charset=a; charset=b` still reads as its media type), and repeated lines are accepted only when they all resolve to the same **supported** format — two agreeing `text/csv` lines are still a `415`. A body declared NDJSON stays NDJSON whatever its bytes, so a bad line is a per-record error rather than a silent re-framing; the reverse (NDJSON sent as `application/json`) is deliberately **not** caught — record one, `200`, the rest ignored ([#561](https://github.com/Wave-RF/WaveHouse/issues/561)). Fail-closed — preserve it when touching `internal/api`. 4. **Async ingestion** — ingest returns 200 after optional dedup + MQ publish; ClickHouse writes happen later via `StartIngestWorker`. NATS full → 503 + Retry-After. 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. -6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format` — a pre-v2 envelope carries none — or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. +6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format`, or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. 8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index ad712b82..9b675b4c 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -612,7 +612,7 @@ Opens a persistent SSE connection for real-time event streaming. Supports histor | ------ | ----------- | | `Last-Event-ID` | RFC 3339 timestamp of the last received event. If present, overrides the `since` query parameter for automatic reconnection (standard `EventSource` behavior). | -**Response:** SSE stream (`text/event-stream`). Data events include an `id:` field set to the event's `received_timestamp`. The stream opens with a `: connected` comment and emits a minimal `:` keepalive comment periodically (every 30 seconds by default), which keeps a quiet connection from being closed by a proxy; both are standard SSE comments that `EventSource` ignores (raw consumers should skip `:`-prefixed lines). When the server stops (see [Stopping](/deployment#stopping)) it ends every open stream immediately rather than holding it for the drain; `EventSource` reconnects on its own and resumes from `Last-Event-ID`. A reload that stops serving the stream's tenant — its folder removed or rejected, over a [nested settings directory](/deployment#the-nested-settings-directory) — ends that tenant's open streams the same way, and the reconnect then gets its `404` (removed) or `503` (rejected): the SDK stops on the `404` and retries the `503`, resuming from `Last-Event-ID` once the folder is back, while a browser `EventSource` treats either as fatal. A browser going cross-origin reads either refusal only when it passes CORS: it is decorated from tenant `0`'s list ([multi-tenant deployments](/deployment#multi-tenant-deployments)), so where tenant `0` is not served or its list does not admit the page's origin, the SDK sees a network error instead and keeps re-dialing. +**Response:** SSE stream (`text/event-stream`). Data events include an `id:` field set to the event's `received_timestamp`. The stream opens with a `: connected` comment and emits a minimal `:` keepalive comment periodically (every 30 seconds by default), which keeps a quiet connection from being closed by a proxy; both are standard SSE comments that `EventSource` ignores (raw consumers should skip `:`-prefixed lines). When the server stops (see [Stopping](/deployment#stopping)) it ends every open stream immediately rather than holding it for the drain; `EventSource` reconnects on its own and resumes from `Last-Event-ID`. A reload that stops serving the stream's tenant — its folder removed or rejected, over a [nested settings directory](/deployment#the-nested-settings-directory) — ends that tenant's open streams the same way, and the reconnect then gets its `404` (removed) or `503` (rejected): the SDK stops on the `404` and retries the `503`, resuming from `Last-Event-ID` once the folder is back — with a hole where the tenant's history was, which the sweeper purges within a minute of the tenant no longer being served — while a browser `EventSource` treats either as fatal. A browser going cross-origin reads either refusal only when it passes CORS: it is decorated from tenant `0`'s list ([multi-tenant deployments](/deployment#multi-tenant-deployments)), so where tenant `0` is not served or its list does not admit the page's origin, the SDK sees a network error instead and keeps re-dialing. **Row values arrive positionally, and the column names are announced separately.** Before the first row, and again whenever the column list changes, the stream sends an `event: schema` frame naming the columns of the rows that follow — in order, already reduced to what the caller's role may read. That re-announcement is **not** guaranteed after a gap-fill across a column change; see the arity note below. Every data frame's `row` array then has exactly one value per announced column, in that order. `schema` is a **named** SSE event, so a browser `EventSource` must `addEventListener('schema', …)` — it never reaches `onmessage`. A schema frame carries **no** `id:` line, so it never moves the client's `Last-Event-ID`. In the example below the table has its own `received_timestamp` **column**, which collides by name with the frame's top-level `received_timestamp` **field** — they are different values: the field is when WaveHouse received the event, the row slot is that column as published (`null` where the record omitted it, which ClickHouse replaces with the column's default on insert). @@ -858,7 +858,7 @@ The message format used on NATS JetStream between ingest and the batch consumer: | `columns` | string[] | The table's **insertable** column names, in declaration order — what each position in `row` means. A `MATERIALIZED` or `ALIAS` column is computed by ClickHouse and cannot be named in an `INSERT`, so it never appears here. | | `row` | array | One `JSONCompactEachRow` line: one value per entry in `columns`, in that order. A column the request body omitted is `null` here; for a **non-nullable** column the insert turns that back into the column's default (`input_format_null_as_default`), but a `Nullable(T) DEFAULT …` column stores `NULL` — only an *absent* key ever took the default, and a positional row has one slot per column and no way to express absence. Parseable `DateTime`/`DateTime64` values are rewritten to canonical RFC 3339 UTC (see [timestamp canonicalization](#timestamp-canonicalization)); other values as originally sent. | -`columns` and `row` are only meaningful together: a reader that cannot pair them — a length mismatch, an undecodable row, a `columns` list naming one column twice — has no way to map a value to a column. Both readers also refuse an envelope whose `format` they do not recognize, which is what a pre-v2 message looks like. Either way the SSE fan-out withholds such an envelope rather than guess, and the batch consumer parks it on the DLQ with `X-DLQ-*` headers — acking and dropping it only where the DLQ is switched off for that table, since it can never insert on retry. Both outcomes increment `wavehouse_ingest_poison_total`, separated by its `disposition` label (`parked` / `dropped`). +`columns` and `row` are only meaningful together: a reader that cannot pair them — a length mismatch, an undecodable row, a `columns` list naming one column twice — has no way to map a value to a column. Both readers also refuse an envelope whose `format` they do not recognize. Either way the SSE fan-out withholds such an envelope rather than guess, and the batch consumer parks it on the DLQ with `X-DLQ-*` headers — acking and dropping it only where the DLQ is switched off for that table, since it can never insert on retry. Both outcomes increment `wavehouse_ingest_poison_total`, separated by its `disposition` label (`parked` / `dropped`). ### Client-Facing Format (SSE) @@ -876,7 +876,7 @@ Three values, where the envelope above has four: this is the frame a role restri ## Dead Letter Queue (DLQ) -When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the tenant's own DLQ NATS stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format`, or `columns` and `row` that do not pair — is parked without ever reaching a table batch. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. +When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the tenant's own DLQ NATS stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format`, or `columns` and `row` that do not pair — is parked without ever reaching a table batch. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: malformed JSON, an envelope of an unknown `format`, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. Use `GET /v1/ops/dlq/stats` to monitor DLQ depth, per tenant (`?tenant=`). diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 92508dac..511a2311 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -148,7 +148,7 @@ The **only** package that imports NATS/JetStream — a `depguard` rule in `.gola - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. It runs on each tenant's stream at that tenant's cutoff. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, or is refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds, plus five more for the rollback (a budget of its own, not the one that just expired), since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, or is refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index b2bac204..996d165f 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -379,7 +379,7 @@ settings/ └── roles.json ``` -That is the layout a control plane writes. Each folder's `clickhouse` block is its tenant's own ClickHouse, so a tenant answers queries once its first schema discovery against that ClickHouse succeeds (until then its schema-aware routes answer `503`, `schema not loaded yet`); what tenant `0`'s folder still supplies for the whole process — the token verifier of the routes that name no tenant, their CORS list — is listed under "What a tenant's folder decides", below. +That is the layout a control plane writes. Each folder's `clickhouse` block is its tenant's own ClickHouse, so a tenant answers queries once its first schema discovery against that ClickHouse succeeds (until then its schema-aware routes answer `503`, `schema not loaded yet`); what tenant `0`'s folder still supplies for the whole process — the token verifier of the routes that name no tenant, their CORS list — is listed under "What a lost tenant `0` costs", below. The folder name is the tenant id, and each folder is a complete settings directory: everything on the [Settings Directory](/settings-directory) page applies to it as written, except where the rules below say otherwise. The two shapes don't mix — a folder beside the four files, or a loose file beside the folders, is a validation error — and a running server keeps the shape it booted with, so switching is stop, restructure, start. The dedupe store needs no restructuring: it keys every tenant's seen ids by tenant, and the four files are tenant `0`, as a `0` folder is. Dot-prefixed entries are ignored in either shape. `wavehouse validate` checks either shape with the same exit codes; a finding in a nested directory names its folder (`acme/policies.json`), and a folder whose name is not a tenant id is a finding of its own — that folder is skipped, and the rest of the directory still loads. @@ -389,7 +389,7 @@ The folder name is the tenant id, and each folder is a complete settings directo **The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. +**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. **What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. @@ -423,7 +423,7 @@ The NATS envelope changed shape in this release: the row now travels positionall This affects the streaming surface too, and more quietly. SSE gap-fill (`?since=` / `Last-Event-ID`) replays from the queue, and the upgrade deletes the old one (next paragraph) — the acknowledged history kept for replay included — so this outlives a *correct* drain: any replay spanning the upgrade silently omits the pre-upgrade events, with **no error and no frame**. Clients that need them should backfill over REST. -**The upgrade does not carry the old queue over at all.** Boot deletes the earlier build's queue and dead-letter queue (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) and everything in them, logging a `WARN` with each one's message count: an event the old build had not yet inserted, and a row it had already parked, do not survive the upgrade. Draining first is the only way to keep them, and anything parked has to be replayed before the upgrade — re-ingest each parked envelope's inner `data` object as a fresh `POST /v1/ingest`. +**The upgrade does not carry the old queue over at all.** Boot deletes the earlier build's queue and dead-letter queue (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) and everything in them, logging a `WARN` with each one's message count: an event the old build had not yet inserted, and a row it had already parked, do not survive the upgrade. Draining first keeps the events not yet inserted; a row already parked is lost with the queue, since the earlier build offers no way to read one back (`GET /v1/ops/dlq/stats` returns counts only). Three audits belong **before** the drain, because none of them announces itself afterwards: @@ -435,10 +435,10 @@ To drain before upgrading: 1. **Stop the producers**, or cut `/v1/ingest` at the reverse proxy. Nothing new should enter the stream. 2. **Wait for the in-flight batches to flush.** A table's batch closes on size or after `maxWait` (5s by default), so a few seconds after the last write is enough; give it longer if ClickHouse is slow or retrying. -3. **Confirm nothing is left unconsumed** before swapping binaries. Not that the stream is empty: it is dual-use, and deliberately retains ACKed messages for the SSE replay window, so a non-zero depth right after a clean drain is expected. Rows landing in ClickHouse is a success signal, **not proof the queue is drained** — when the DLQ is off for a table, a row that fails its retry is skipped without being acked, so NATS keeps redelivering it while its neighbors land. Check that nothing is still failing or redelivering, and note which signal covers which case: [`GET /v1/ops/dlq/stats`](/api#get-v1opsdlqstats--dlq-statistics) is non-zero only where the DLQ is **on**; `wavehouse_ingest_poison_total` counts unreadable envelopes on **either** setting, told apart by its `disposition` label (`parked` / `dropped`); and for a twice-failed row with the DLQ off — the case just described — the **only** signal is the `ERROR` log (`isolated bad row, DLQ disabled for table`). A clean `dlq/stats` with the DLQ off proves nothing. There is no queue-depth gauge today ([#544](https://github.com/Wave-RF/WaveHouse/issues/544) tracks the related in-flight accounting), and `wavehouse_nats_in_msgs_total` going flat is a supporting signal rather than a guarantee. Enabling the DLQ is not itself a drain — replay from `dlq.{table}` is manual, and has to happen before the upgrade deletes it. +3. **Confirm nothing is left unconsumed** before swapping binaries. Not that the stream is empty: it is dual-use, and deliberately retains ACKed messages for the SSE replay window, so a non-zero depth right after a clean drain is expected. Rows landing in ClickHouse is a success signal, **not proof the queue is drained** — when the DLQ is off for a table, a row that fails its retry is skipped without being acked, so NATS keeps redelivering it while its neighbors land. Check that nothing is still failing or redelivering, and note which signal covers which case: [`GET /v1/ops/dlq/stats`](/api#get-v1opsdlqstats--dlq-statistics) is non-zero only where the DLQ is **on**; `wavehouse_ingest_poison_total` counts unreadable envelopes on **either** setting, told apart by its `disposition` label (`parked` / `dropped`); and for a twice-failed row with the DLQ off — the case just described — the **only** signal is the `ERROR` log (`isolated bad row, DLQ disabled for table`). A clean `dlq/stats` with the DLQ off proves nothing. There is no queue-depth gauge today ([#544](https://github.com/Wave-RF/WaveHouse/issues/544) tracks the related in-flight accounting), and `wavehouse_nats_in_msgs_total` going flat is a supporting signal rather than a guarantee. Enabling the DLQ is not itself a drain: a parked row is not inserted, and the upgrade deletes it. 4. **Upgrade**, then re-enable ingest. -If you skipped the drain, the boot's `WARN` line for each deleted stream (`deleted the stream an earlier build kept for every tenant together`) says how many messages went with it. +If you skipped the drain, the boot's `WARN` line for each deleted stream (`deleted the stream an earlier build kept for every tenant together`) says how many messages went with it: for `WAVEHOUSE_DLQ`, the parked rows lost; for `WAVEHOUSE`, a count that includes the acknowledged history kept for replay, already in ClickHouse — so it bounds the events lost rather than counting them, and is non-zero even after a clean drain. ## Dead Letter Queue (DLQ) diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index b504fef4..72447024 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -22,7 +22,7 @@ The pipeline is **insert-only**. (Upgrading across the v2 envelope? [Drain the q ## High-level shape -Each tenant's events are queued on a JetStream stream of its own. One process holds one durable consumer on each tenant's stream, delivered into one handler, and fans events out to a goroutine per tenant table — the tenant is the subject's leading token. Each tenant's table batches independently and POSTs to ClickHouse over the HTTP interface (`JSONCompactEachRow`). On a bulk-insert failure the batch is re-inserted row by row, so a single poison row can't sink it: clean rows ack, and only the rows that fail again go to the dead-letter stream. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and meets the dead-letter switch once, whole; a tenant no longer served has no switch to read, so its batch is parked. An envelope the worker cannot *read* — malformed JSON, an unknown row `format` (what a pre-v2 message looks like), or columns and a row that don't pair — never reaches a table loop at all: `parseMsg` parks it on the same dead-letter stream, or, where the DLQ is off for the table, acks and drops it rather than redelivering a message that can never insert. A separate sweeper reclaims stream storage. +Each tenant's events are queued on a JetStream stream of its own. One process holds one durable consumer on each tenant's stream, delivered into one handler, and fans events out to a goroutine per tenant table — the tenant is the subject's leading token. Each tenant's table batches independently and POSTs to ClickHouse over the HTTP interface (`JSONCompactEachRow`). On a bulk-insert failure the batch is re-inserted row by row, so a single poison row can't sink it: clean rows ack, and only the rows that fail again go to the dead-letter stream. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and meets the dead-letter switch once, whole; a tenant no longer served has no switch to read, so its batch is parked. An envelope the worker cannot *read* — malformed JSON, an unknown row `format`, or columns and a row that don't pair — never reaches a table loop at all: `parseMsg` parks it on the same dead-letter stream, or, where the DLQ is off for the table, acks and drops it rather than redelivering a message that can never insert. A separate sweeper reclaims stream storage. ```mermaid flowchart LR diff --git a/docs/src/content/docs/sdk/streaming.md b/docs/src/content/docs/sdk/streaming.md index 94a7e334..2b37667a 100644 --- a/docs/src/content/docs/sdk/streaming.md +++ b/docs/src/content/docs/sdk/streaming.md @@ -116,7 +116,7 @@ A dropped stream reconnects on a jittered exponential backoff, capped at 30s, an :::caution[Resumption is at-least-once, and time-bounded] Delivery across a reconnect is **at-least-once**. The `Last-Event-ID` the client sends is the last event's `received_timestamp`, and the server replays from that instant *inclusively* — so the last event you already saw, and anything sharing its timestamp, arrives again. The SDK does not deduplicate live frames — `liveQuery()` makes one pass at the backfill seam, and only under an ascending order ([#449](https://github.com/Wave-RF/WaveHouse/issues/449)) — so key on `timestamp` plus your own row identity if duplicates matter. -Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. The same silence applies across a server upgrade to this release: the server deletes the previous release's queue at boot, so a replay spanning the upgrade omits the events published before it, without an error — backfill over REST if you need them. +Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. So does a stream a [nested server](/deployment#the-nested-settings-directory) ended because its tenant's folder was rejected, once the folder is fixed: the sweeper purges a tenant's history within a minute of the tenant no longer being served. The same silence applies across a server upgrade to this release: the server deletes the previous release's queue at boot, so a replay spanning the upgrade omits the events published before it, without an error — backfill over REST if you need them. **A column-set change across a gap-fill is a known limitation.** If the table's columns change while you are connected *and* your client replays across that change, live rows arriving after the replay may not be preceded by a fresh `event: schema` frame until the columns next change or you reconnect. The SDK drops a row whose **length** disagrees with the list it was last told, rather than zipping it under the wrong names — so an added or removed column costs you rows, not wrong ones. A **same-length** change is the residual case the arity check cannot see: a `RENAME COLUMN`, or a drop paired with an add, zips values under the wrong names until the next announcement. Reconnecting resynchronizes either way. Full schema-change handling is deferred to the schema-versioning work ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)). ::: diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 41e04187..29f842bc 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -213,7 +213,7 @@ The `auth` block is the verifier wiring, minus the secrets. `jwks_url` (absolute A failed batch insert is retried row by row; a row that fails again on its own is a poison row. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the retry, which no row of it could pass, and every row of it is a poison row. `dlq.enabled` (seed default `true`) decides what happens to it, resolved per table (`dlq.tables.
.enabled` → global) at the moment of the failure, so a reload applies to the next poison row: - `true` — the row is published to the tenant's dead-letter stream (`DLQ_{tenant}`) under `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) with the failure in its headers, and its original is acked. Inspect it with `GET /v1/ops/dlq/stats` (admin-only; `?tenant=` names the tenant). -- `false` — the row is left unacked, so NATS redelivers it and it retries until it inserts or the switch is flipped back. For every row the worker **can read**, nothing is ever dropped either way — the choice is *park it* versus *keep retrying*. **One exception, new in this release:** an envelope the worker cannot read *at all* — malformed JSON, an unknown `format`, or `columns` and `row` that do not pair — can never insert, so redelivering it forever would wedge the consumer. With the DLQ off for the table it is acked and **dropped**, logged at `ERROR` and counted by `wavehouse_ingest_poison_total` with `disposition="dropped"` (also labeled by `table` and `reason`; an envelope parked on the DLQ carries `disposition="parked"`). See [Ingest Pipeline](/ingest-pipeline) — and drain the ingest queue before upgrading. +- `false` — the row is left unacked, so NATS redelivers it and it retries until it inserts or the switch is flipped back. For every row the worker **can read**, nothing is ever dropped either way — the choice is *park it* versus *keep retrying*. **One exception, new in this release:** an envelope the worker cannot read *at all* — malformed JSON, an unknown `format`, or `columns` and `row` that do not pair — can never insert, so redelivering it forever would wedge the consumer. With the DLQ off for the table it is acked and **dropped**, logged at `ERROR` and counted by `wavehouse_ingest_poison_total` with `disposition="dropped"` (also labeled by `table` and `reason`; an envelope parked on the DLQ carries `disposition="parked"`). See [Ingest Pipeline](/ingest-pipeline). For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in its ingest queue, redelivered for as long as the tenant is away and stopping the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging that queue. diff --git a/internal/app/app.go b/internal/app/app.go index dc15bf5d..6d39d0fe 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -144,7 +144,8 @@ const ( ) // New wires every component. ctx bounds construction only — the boot-time -// schema refresh and the JetStream stream setup; the loops start in Run. A +// schema refresh and the opening of each served tenant's queue; the loops +// start in Run. A // failure releases whatever was already opened and returns the error, so // the caller never holds a half-built App. func New(ctx context.Context, opts Options) (app *App, err error) { @@ -177,7 +178,7 @@ func New(ctx context.Context, opts Options) (app *App, err error) { if err := a.wireDedupe(); err != nil { return nil, err } - if err := a.wireMQ(); err != nil { + if err := a.wireMQ(ctx); err != nil { return nil, err } if err := a.wireCache(); err != nil { diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d9..cda01e6f 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -571,6 +571,50 @@ func TestNew_DedupeOpenFailure(t *testing.T) { }) } +// A tenant's queue the MQ cannot open follows the registry's rule for the +// shape, as the dedupe store does: a flat directory refuses boot, and a nested +// one boots with that tenant's queue closed and every other tenant's open. +// The obstacle is a regular file where the embedded server keeps a stream's +// store — the embedded implementation's layout, which this test takes on to +// force the failure, as TestNew_DedupeOpenFailure does Pebble's. The failed +// open clears it, so the next publish opens the queue: each one tries again. +func TestNew_QueueOpenFailure(t *testing.T) { + block := func(t *testing.T, dataDir, stream string) { + t.Helper() + p := filepath.Join(dataDir, "nats", "jetstream", "$G", "streams", stream) + require.NoError(t, os.MkdirAll(filepath.Dir(p), 0o750)) + require.NoError(t, os.WriteFile(p, nil, 0o600)) + } + t.Run("flat refuses boot", func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, nil)) + block(t, cfg.DataDir, "DLQ_0") + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "mq open") + }) + t.Run("nested costs the tenant alone", func(t *testing.T) { + cfg := testConfig(t, writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil})) + block(t, cfg.DataDir, "DLQ_acme") + a := newApp(t, cfg, Options{}) + assert.Zero(t, a.mq.MaxBytes("acme"), "acme's queue did not open") + assert.Equal(t, int64(50<<30), a.mq.MaxBytes("globex"), "and costs globex nothing") + + require.NoError(t, a.MQ().Publish(t.Context(), mq.Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(50<<30), a.mq.MaxBytes("acme"), "a publish opened it at acme's budget") + }) +} + +// Boot opens each served tenant's queue under New's context, as New's doc +// says: a stop signalled during boot is not held up by one open per tenant. +func TestNew_QueueSetupHonorsTheBootContext(t *testing.T) { + guardGlobals(t) + ctx, cancel := context.WithCancel(t.Context()) + cancel() + _, err := New(ctx, Options{Config: testConfig(t, writeSettings(t, nil))}) + require.ErrorIs(t, err, context.Canceled) + require.ErrorContains(t, err, "mq open") +} + // The tenants on the writer's ClickHouse address and database read the same // tables, so an insert invalidates a table's cached results under every one // of them — whatever their user, so across pools — and under no tenant on diff --git a/internal/app/wire.go b/internal/app/wire.go index 60cbdcd8..565e6bc2 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -541,9 +541,11 @@ func (a *App) wireDedupe() error { // resized follows the registry's rule for the shape: a flat directory // refuses boot, like every other store, and on a reload logs it, keeping the // previous budget; a nested directory logs it at boot too, so it never costs -// the process — the tenant's ingest answers 503 until a reload opens its -// queue. The hook is registered before the boot apply, as the dedupe one is. -func (a *App) wireMQ() error { +// the process — the tenant's ingest answers 503 until its queue opens, each +// publish and each reload trying again. The hook is registered before the +// boot apply, as the dedupe one is. The boot apply runs on ctx, New's, so a +// stop signalled during a boot that opens many queues is not held up by them. +func (a *App) wireMQ(ctx context.Context) error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker @@ -565,17 +567,21 @@ func (a *App) wireMQ() error { } } - // Rooted in the App's stop context, so a reload caught mid-hook by - // SIGTERM gives up rather than holding the drain past - // server.shutdown_timeout. - reconcile := func() error { + // The hook's apply is rooted in the App's stop context, so a reload + // caught mid-hook by SIGTERM gives up rather than holding the drain past + // server.shutdown_timeout; a done ctx ends the pass over the tenants. + reconcile := func(ctx context.Context) error { var errs []error for id, store := range a.tenants.All() { + if err := ctx.Err(); err != nil { + errs = append(errs, err) + break + } mb := store.MQMaxBytes() if mb == broker.MaxBytes(id) { continue } - if err := broker.SetMaxBytes(a.stopCtx, id, mb); err != nil { + if err := broker.SetMaxBytes(ctx, id, mb); err != nil { slog.Error("mq queue not reconciled with settings; the next reload retries", "tenant", id, "error", err) errs = append(errs, fmt.Errorf("tenant %s: %w", id, err)) continue @@ -584,8 +590,8 @@ func (a *App) wireMQ() error { } return errors.Join(errs...) } - a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile() }) - if err := reconcile(); err != nil && !a.tenants.Nested() { + a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) + if err := reconcile(ctx); err != nil && !a.tenants.Nested() { return fmt.Errorf("mq open: %w", err) } return nil diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index e082915c..8be913c5 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -96,7 +96,9 @@ const ( // fails: in-process JetStream fails by stalling rather than erroring, so // the likely cause is that resizeTimeout has just run out, and an undo on // that context would fail without touching the stream. SetMaxBytes runs - // for at most the sum of the two. + // for at most the sum of the two when it resizes, and for two + // resizeTimeouts when it opens a queue: the consumers join on a budget of + // their own (apply). rollbackTimeout = 5 * time.Second ) @@ -303,7 +305,8 @@ func (e *EmbeddedNATS) MaxBytes(id tenant.ID) int64 { // call with the new budget reapplies both. // // The JetStream calls are bounded by resizeTimeout, plus rollbackTimeout for -// the undo, both rooted in ctx. That is deliberate: ctx is the process's stop +// the undo — or another resizeTimeout for the consumers joining a queue just +// opened — all rooted in ctx. That is deliberate: ctx is the process's stop // context, so a reload caught mid-hook by a stop gives up — undo included — // rather than holding the drain past server.shutdown_timeout. A cancellation // between the two updates is therefore the one way to leave the pair split, From 999c1db551d8cd34a63a6ca0d019bcfab7f80709 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 19:29:53 -0400 Subject: [PATCH 004/108] fix(mq): boot re-applies a budget to a split queue pair; review fixes --- docs/src/content/docs/architecture.md | 5 ++- docs/src/content/docs/deployment.md | 4 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app.go | 5 +-- internal/app/app_test.go | 2 +- internal/app/wire.go | 2 +- internal/ingest/worker.go | 9 ++--- internal/mq/embedded.go | 30 ++++++++++++--- internal/mq/embedded_test.go | 40 ++++++++++++++++++++ internal/mq/mq.go | 24 ++++++------ 10 files changed, 92 insertions(+), 31 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 511a2311..549622f6 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -231,7 +231,8 @@ Client POST /v1/ingest?table={table} → (If the tenant's NATS stream is full, or not open: 503 + Retry-After header) Ingest worker pipeline (StartIngestWorker): - ← JetStream pull consumer (buffer-consumer) on ingest.> + ← JetStream pull consumer (buffer-consumer), one durable per tenant stream + (ingest.{tenant}.>), delivered into one handler → Parse the event envelope (an envelope the worker cannot read — malformed JSON, an unknown or absent format, columns and row that don't pair — is parked on the DLQ, or acked-and-dropped where the DLQ is off for the table; either way diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 996d165f..50bedaaa 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -419,9 +419,9 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres ## Upgrading across the v2 ingest envelope -The NATS envelope changed shape in this release: the row now travels positionally, with `format`, `columns` and `row` replacing `data`. **The new worker cannot read a message published by an older version** — it carries no `format`, so there is no way to say which value belongs to which column. +The NATS envelope changed shape in this release: the row now travels positionally, with `format`, `columns` and `row` replacing `data` — and the queue changed layout with it: boot deletes the earlier build's queue (below), so nothing an older version published reaches the new worker, which could not read it anyway (it carries no `format`, so there is no way to say which value belongs to which column). **Drain first** to keep what the old build had not yet inserted. -This affects the streaming surface too, and more quietly. SSE gap-fill (`?since=` / `Last-Event-ID`) replays from the queue, and the upgrade deletes the old one (next paragraph) — the acknowledged history kept for replay included — so this outlives a *correct* drain: any replay spanning the upgrade silently omits the pre-upgrade events, with **no error and no frame**. Clients that need them should backfill over REST. +The streaming surface loses something too, more quietly. SSE gap-fill (`?since=` / `Last-Event-ID`) replays from the queue, so the deletion takes the replay history with it, even after a *correct* drain: any replay spanning the upgrade silently omits the pre-upgrade events, with **no error and no frame**. Clients that need them should backfill over REST. **The upgrade does not carry the old queue over at all.** Boot deletes the earlier build's queue and dead-letter queue (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) and everything in them, logging a `WARN` with each one's message count: an event the old build had not yet inserted, and a row it had already parked, do not survive the upgrade. Draining first keeps the events not yet inserted; a row already parked is lost with the queue, since the earlier build offers no way to read one back (`GET /v1/ops/dlq/stats` returns counts only). diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 29f842bc..a0808514 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -221,7 +221,7 @@ A tenant's dead-letter stream exists from the moment the tenant is first served ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume. A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume. A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. ## Streaming diff --git a/internal/app/app.go b/internal/app/app.go index 6d39d0fe..2131942c 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -145,9 +145,8 @@ const ( // New wires every component. ctx bounds construction only — the boot-time // schema refresh and the opening of each served tenant's queue; the loops -// start in Run. A -// failure releases whatever was already opened and returns the error, so -// the caller never holds a half-built App. +// start in Run. A failure releases whatever was already opened and returns the +// error, so the caller never holds a half-built App. func New(ctx context.Context, opts Options) (app *App, err error) { a := &App{cfg: opts.Config, build: opts.Build, logLevel: opts.LogLevel, listener: opts.Listener} if a.logLevel == nil { diff --git a/internal/app/app_test.go b/internal/app/app_test.go index cda01e6f..8ef1f7c5 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -605,7 +605,7 @@ func TestNew_QueueOpenFailure(t *testing.T) { } // Boot opens each served tenant's queue under New's context, as New's doc -// says: a stop signalled during boot is not held up by one open per tenant. +// says: a stop signaled during boot is not held up by one open per tenant. func TestNew_QueueSetupHonorsTheBootContext(t *testing.T) { guardGlobals(t) ctx, cancel := context.WithCancel(t.Context()) diff --git a/internal/app/wire.go b/internal/app/wire.go index 565e6bc2..e08f4a0e 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -544,7 +544,7 @@ func (a *App) wireDedupe() error { // the process — the tenant's ingest answers 503 until its queue opens, each // publish and each reload trying again. The hook is registered before the // boot apply, as the dedupe one is. The boot apply runs on ctx, New's, so a -// stop signalled during a boot that opens many queues is not held up by them. +// stop signaled during a boot that opens many queues is not held up by them. func (a *App) wireMQ(ctx context.Context) error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index b4defa36..f6cbaf6a 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -226,11 +226,10 @@ func waitOrDeadline(ctx context.Context, wg *sync.WaitGroup) error { // dispatchLoop owns the one consumer — held on every tenant's queue — and fans // every message out to a tableLoop per tenant table (lazily spawned on first -// sight of one). It does -// no batching itself — it parses just enough to route — so a low-volume table -// can never strand another table's rows behind a shared timer. It is the ONLY -// goroutine that watches ctx; tableLoops stop via channel-close, which gives a -// deterministic drain with no abandoned messages. +// sight of one). It does no batching itself — it parses just enough to route — +// so a low-volume table can never strand another table's rows behind a shared +// timer. It is the ONLY goroutine that watches ctx; tableLoops stop via +// channel-close, which gives a deterministic drain with no abandoned messages. func (w *IngestWorker) dispatchLoop(ctx context.Context, cons mq.Consumer) { defer w.wg.Done() diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 8be913c5..b621f96e 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -74,8 +74,9 @@ type tenantQueue struct { ingest, dlq bool // maxBytes is the budget last applied in full (MaxBytes); asked is the // budget last asked for, which a publish or park that finds a stream - // missing opens it at. Both are read back from the ingest stream at boot, - // so a tenant no longer served keeps the budget it last had. + // missing opens it at. Boot reads asked back from the ingest stream, so a + // tenant no longer served keeps the budget it last had, and maxBytes too + // when the pair is whole at it (takeStock). maxBytes, asked int64 } @@ -182,20 +183,39 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { return err } } + type dlqState struct { + limit int64 + held uint64 + } + dlqs := map[tenant.ID]dlqState{} streams := e.js.ListStreams(ctx) for info := range streams.Info() { name := info.Config.Name if id, ok := streamTenant(ingestStreamPrefix, name); ok { q := e.queue(id) q.ingest = true - q.maxBytes, q.asked = info.Config.MaxBytes, info.Config.MaxBytes + q.asked = info.Config.MaxBytes } else if id, ok := streamTenant(dlqStreamPrefix, name); ok { e.queue(id).dlq = true + dlqs[id] = dlqState{limit: info.Config.MaxBytes, held: info.State.Bytes} } } if err := streams.Err(); err != nil { return fmt.Errorf("list streams: %w", err) } + // A pair is at its budget when its dead-letter stream is at a tenth of + // the ingest cap, or above it holding more than that: the shrink guard's + // doing. Anything else is a pair a stop or a failed update left split, or + // one missing its dead-letter stream, so its budget stays unapplied and + // the boot's SetMaxBytes applies it to both streams again. + for id, q := range e.queues { + d, ok := dlqs[id] + tenth := q.asked / dlqShare + guarded := d.limit > tenth && d.held <= math.MaxInt64 && int64(d.held) > tenth + if q.ingest && ok && (d.limit == tenth || guarded) { + q.maxBytes = q.asked + } + } return nil } @@ -668,8 +688,8 @@ func sameConsumer(have, want jetstream.ConsumerConfig) bool { } // share is one tenant's part of the fetch-ahead: the total spread over the -// tenants' queues, at least one each. 0 leaves the client default. Under -// e.mu. +// tenants' queues joined so far, at least one each, fixed when that queue's +// delivery starts. 0 leaves the client default. Under e.mu. func (f *fanIn) share() int { if f.prefetch <= 0 { return 0 diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 06120cc5..e75b62d7 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -1083,6 +1083,46 @@ func TestNewEmbedded_DeletesTheStreamsAnEarlierBuildShared(t *testing.T) { require.NoError(t, e.Publish(ctx, Topic{Tenant: tenant.Default, Table: "events"}, []byte("x"))) } +// A pair a stop or a failed update left split — its dead-letter stream not +// at a tenth of the ingest cap — or one missing its dead-letter stream is not +// at its budget, so the boot's SetMaxBytes applies the budget to both streams +// again; a dead-letter stream kept above its tenth because it holds more (the +// shrink guard) is at its budget and left as it is. +func TestNewEmbedded_ASplitPairIsAppliedAgainAtBoot(t *testing.T) { + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + dir := t.TempDir() + first, err := NewEmbedded(dir) + require.NoError(t, err) + for _, id := range []tenant.ID{"split", "gone", "guarded"} { + require.NoError(t, first.SetMaxBytes(ctx, id, 10<<20)) + } + _, err = first.js.UpdateStream(ctx, dlqStreamConfig("split", 2<<20)) + require.NoError(t, err) + require.NoError(t, first.js.DeleteStream(ctx, "DLQ_gone")) + payload := make([]byte, 1<<10) + for range 200 { + require.NoError(t, first.DeadLetter(ctx, NewMessage(ctx, Topic{Tenant: "guarded", Table: "t"}, payload, time.Now(), nil, nil, nil))) + } + require.NoError(t, first.SetMaxBytes(ctx, "guarded", 1<<20)) + guardedCap := streamConfig(t, first, "DLQ_guarded").MaxBytes + require.Greater(t, guardedCap, int64(1<<20)/10, "the guard kept the parked rows") + require.NoError(t, first.Close()) + + e := openEmbedded(t, dir) + assert.Zero(t, e.MaxBytes("split"), "a split pair is not at its budget") + assert.Zero(t, e.MaxBytes("gone"), "nor one missing its dead-letter stream") + assert.Equal(t, int64(1<<20), e.MaxBytes("guarded"), "a guarded dead-letter stream is") + + for _, id := range []tenant.ID{"split", "gone"} { + require.NoError(t, e.SetMaxBytes(ctx, id, 10<<20)) + assert.Equal(t, int64(10<<20), e.MaxBytes(id)) + assert.Equal(t, int64(10<<20)/10, streamConfig(t, e, dlqStreamName(id)).MaxBytes, "%s: the pair is whole again", id) + } + require.NoError(t, e.SetMaxBytes(ctx, "guarded", 1<<20)) + assert.Equal(t, guardedCap, streamConfig(t, e, "DLQ_guarded").MaxBytes, "left as the guard kept it") +} + // A boot takes stock of the queues on disk: each keeps the budget it last // had, and a consumer created afterwards is held on every one of them — a // tenant no longer served, which is never given a budget again, included — diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 34620f85..663bde3d 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -195,17 +195,19 @@ type ConsumerConfig struct { // Consumer is a live durable consumer created by ConsumerManager. type Consumer interface { - // Consume delivers each message to handler on a delivery goroutine of - // its tenant's: one per tenant, so a tenant's messages arrive in order, - // one at a time, while different tenants' arrive concurrently — handler - // must be safe for that. A handler that blocks holds back its tenant's - // delivery — that is the backpressure the ingest worker relies on. About - // prefetch messages are fetched ahead across the tenants together, at - // least one per tenant (0 = the client default, per tenant). The returned - // stop asks delivery to end and returns without waiting: a handler - // invocation already in flight, or one for a message already queued - // client-side, may still run after stop returns, so a handler must not - // write to anything the caller tears down right after stopping. + // Consume delivers each message to handler on a delivery goroutine of its + // tenant's: one per tenant, so a tenant's messages arrive in order, one at + // a time, while different tenants' arrive concurrently — handler must be + // safe for that. A handler that blocks holds back its tenant's delivery — + // that is the backpressure the ingest worker relies on. About prefetch + // messages are fetched ahead across the tenants together: the tenants' + // queues when delivery starts split it, and a queue joined later fetches + // ahead its share of it at that point, at least one message each (0 = the + // client default, per tenant). The returned stop asks delivery to end and + // returns without waiting: a handler invocation already in flight, or one + // for a message already queued client-side, may still run after stop + // returns, so a handler must not write to anything the caller tears down + // right after stopping. // // Delivery can also end on its own after Consume has returned: the broker // or the client gives up on the consumer (it was deleted, the connection From db2d20a20a4f3af04ac671d1fcc2e7387215cd4b Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 20:01:33 -0400 Subject: [PATCH 005/108] fix(mq): a failed resize restores the ingest stream's own cap; review fixes --- CHANGELOG.md | 2 +- docs/src/content/docs/deployment.md | 6 ++-- docs/src/content/docs/durability.md | 6 ++-- internal/mq/embedded.go | 21 ++++++++----- internal/mq/embedded_test.go | 49 +++++++++++++++++++++++++++++ 5 files changed, 70 insertions(+), 14 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c715..86dac970 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -24,7 +24,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Schema discovery captures each table's DDL, its columns' ordinals and default expressions, and the server version** (`internal/discovery/discovery.go`, `internal/testutil/testutil.go`): `Column` gains `DefaultExpression` and `Position` (both from a widened `system.columns` select), `TableSchema` gains `DDL` from `system.tables.create_table_query`, and `SchemaRegistry` gains `ServerVersion()` from a `SELECT version()` probe next to the existing `SELECT timezone()`. Groundwork for the native type layer, captured on the same refresh as the columns so a stale version cannot outlive the schemas it describes. That is a publication guarantee, not a same-server one: `chconn.Manager` resolves the connection per call, so a reload changing `clickhouse.addr` mid-refresh can still pair a version from one server with schemas from another — narrow, and self-correcting on the next refresh. `DDL` is `json:"-"` and does **not** appear in `/v1/ops/schema`: that endpoint marshals `TableSchema` straight to the client, and an external-engine table (S3, MySQL, PostgreSQL, Kafka) renders its wiring there unconditionally — endpoint, bucket or host, database, username, S3 access key id. ClickHouse masks the password itself as `[HIDDEN]` from ~23.9 (verified on 26.7.3), so the exposure is the topology rather than the secret — except on an older server, or one with `display_secrets_in_show_and_select` enabled. `position` and `default_expression` are additive fields in the response. A table listed in `system.tables` with no `system.columns` rows is skipped rather than published column-less, and both new queries fail the refresh on error exactly as `timezone()` and `system.columns` do — callers keep the prior cache and retry. -- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the worker drains, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. +- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the sweeper purges it back under the limit, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. - **"Was this page helpful?" feedback widget on every docs page** (`docs/src/components/PageFeedback.astro` (new), `docs/src/components/Footer.astro`): a thumbs-up / thumbs-down vote below the page content, captured to PostHog as `docs_feedback` with `{ helpful, page }`. It renders from `Footer.astro`'s sidebar branch — the same indirection the Cloud CTA uses — rather than a per-page import or frontmatter flag, so every content page gets it automatically, including ones not written yet; it sits *below* the Cloud CTA on the pages that carry one, and splash pages (the homepage and 404) take the other footer branch and never render it. One vote per page per visitor: the choice is remembered in `localStorage` keyed by pathname, and a revisit renders the thanks message instead of re-prompting (storage is a nicety, not the record — a browser with storage disabled still votes). - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 50bedaaa..2029bc0d 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -383,13 +383,13 @@ That is the layout a control plane writes. Each folder's `clickhouse` block is i The folder name is the tenant id, and each folder is a complete settings directory: everything on the [Settings Directory](/settings-directory) page applies to it as written, except where the rules below say otherwise. The two shapes don't mix — a folder beside the four files, or a loose file beside the folders, is a validation error — and a running server keeps the shape it booted with, so switching is stop, restructure, start. The dedupe store needs no restructuring: it keys every tenant's seen ids by tenant, and the four files are tenant `0`, as a `0` folder is. Dot-prefixed entries are ignored in either shape. `wavehouse validate` checks either shape with the same exit codes; a finding in a nested directory names its folder (`acme/policies.json`), and a folder whose name is not a tenant id is a finding of its own — that folder is skipped, and the rest of the directory still loads. -**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). Its message queue is kept, at the budget it last had, but the history gap-fill replays is purged from it at the next sweep, as a removed tenant's is, so a stream resumed after the fix has a hole where that history was. The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. +**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). Its message queue is kept, at the budget it last had, but the history that gap-fill replays is purged from it at the next sweep, as a removed tenant's is, so a stream resumed after the fix has a hole where that history was. The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. -**Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted — its message queue is kept at the budget it last had, and only the history gap-fill replays goes from it, at the next sweep — so restoring the folder restores the tenant, seen ids and parked rows included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. +**Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted — its message queue is kept at the budget it last had, and only the history that gap-fill replays goes from it, at the next sweep — so restoring the folder restores the tenant, seen ids and parked rows included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. **The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. +**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history that gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. **What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 353608f2..4f7cc1b7 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -33,7 +33,7 @@ WaveHouse does not currently expose a knob to relax this — `SyncAlways` is alw Because the publish blocks on `fsync`, **your typical ingest latency is your storage's typical `fsync` latency, and your worst-case publish is your storage's worst-case `fsync`.** When that tail is healthy (sub-millisecond to single-digit milliseconds) the guarantee is essentially free. When it is not, the same code path that handles every production message stalls: - Publishes block for the duration of the `fsync`, so a multi-second `fsync` tail is a multi-second ingest tail. -- The embedded server's stream/consumer setup and every publish run under the JetStream client's request timeout; a slow-enough substrate makes them exceed it. The symptom at a first boot, which opens every tenant's queue, is `open dlq stream: ... context deadline exceeded`; a later boot writes nothing, so the first publish is where it shows. +- The embedded server's consumer setup and every publish run under the JetStream client's request timeout, and opening or resizing a tenant's queue under a ten-second budget of WaveHouse's own; a slow-enough substrate makes them exceed it. The symptom when a tenant's queue first opens — at the boot or reload that first serves the tenant — is `open dlq stream: ... context deadline exceeded`, or `open ingest stream: ...` (the two share the budget); a boot that finds every queue already at its budget writes nothing, so there the first publish is where it shows. - If the worker cannot drain to ClickHouse faster than producers publish, a tenant's stream fills toward its [`mq.max_bytes_gb`](/settings-directory#message-queue) and the API returns `503` to that tenant ([backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs)). ## Where `SyncAlways` is cheap vs. expensive @@ -81,7 +81,7 @@ Read the measured p99 against these bands, which track WaveHouse's `SyncAlways` | 1–5 ms | **Good** | | 5–50 ms | **Workable** — watch bursty load | | 50 ms – 1 s | **Marginal** — relax durability once `mq.sync_interval` ([#139](https://github.com/Wave-RF/WaveHouse/issues/139)) lands, or move to faster storage | -| > 1 s | **Broken** — opening a tenant's queue (`open dlq stream`) will time out under load; fix the storage substrate | +| > 1 s | **Broken** — opening a tenant's queue (`open dlq stream` / `open ingest stream`) will time out under load; fix the storage substrate | :::caution[macOS `fsync` lies by default] A plain `fsync()` on macOS returns once data is in the drive's volatile cache — it does **not** force a flush to NAND; only `fcntl(fd, F_FULLFSYNC)` does (NATS, Postgres, and SQLite all use it). On a Mac, any per-flush number under ~1 ms is almost certainly not a real flush — the gap between plain `fsync()` and `F_FULLFSYNC` can be ~180× on the same consumer NVMe. `fio` on macOS calls plain `fsync()`, so don't trust Mac `fio` numbers for tail-latency planning. This mostly matters when benchmarking a dev machine; production WaveHouse runs on Linux, where `fio` is honest. @@ -93,7 +93,7 @@ A self-contained `wavehouse storage-check` preflight subcommand that bakes this If you see any of these, benchmark the `/nats` volume as above: -- `open dlq stream: ... context deadline exceeded` when a tenant's queue first opens, at the first boot or at the reload that adopts the tenant. +- `open dlq stream: ... context deadline exceeded`, or `open ingest stream: ...`, when a tenant's queue first opens, at the boot or reload that first serves the tenant. - Ingest p99 latency in the seconds, or occasional `200`s that take multiple seconds to return. - Intermittent `503 Service Unavailable` from `/v1/ingest` when ClickHouse is healthy (the worker can't drain fast enough because acking is `fsync`-bound). - Flaky CI or load tests that pass on fast storage and fail on a shared/virtualized host. diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index b621f96e..c3762a8f 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -78,6 +78,10 @@ type tenantQueue struct { // tenant no longer served keeps the budget it last had, and maxBytes too // when the pair is whole at it (takeStock). maxBytes, asked int64 + // ingestCap is the cap the ingest stream has — what a failed resize + // restores it to. Not maxBytes: a pair boot found split has a cap but no + // budget applied in full, and a cap of 0 would be none at all. + ingestCap int64 } // EmbeddedNATS is the one implementation of every mq interface. @@ -194,7 +198,7 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { if id, ok := streamTenant(ingestStreamPrefix, name); ok { q := e.queue(id) q.ingest = true - q.asked = info.Config.MaxBytes + q.asked, q.ingestCap = info.Config.MaxBytes, info.Config.MaxBytes } else if id, ok := streamTenant(dlqStreamPrefix, name); ok { e.queue(id).dlq = true dlqs[id] = dlqState{limit: info.Config.MaxBytes, held: info.State.Bytes} @@ -310,10 +314,10 @@ func (e *EmbeddedNATS) MaxBytes(id tenant.ID) int64 { // JetStream applies a limit change to a live stream without touching its // messages: growing takes effect immediately; shrinking the ingest stream // below its current size makes DiscardNew refuse new publishes until the -// worker drains it — nothing buffered is dropped. The dead-letter stream is -// DiscardOld, which would delete its oldest parked rows to fit a smaller cap, -// so it is never capped below the bytes it holds (#532): it keeps what it -// has, and that is logged. +// sweeper purges it back under the cap — nothing buffered is dropped. The +// dead-letter stream is DiscardOld, which would delete its oldest parked rows +// to fit a smaller cap, so it is never capped below the bytes it holds (#532): +// it keeps what it has, and that is logged. // // The pair moves together where it can. If the dead-letter update fails after // the ingest one succeeded, the ingest resize is undone so the pair stays at @@ -358,7 +362,7 @@ func (e *EmbeddedNATS) apply(ctx context.Context, id tenant.ID, q *tenantQueue, if _, err := e.js.CreateOrUpdateStream(resizeCtx, ingestStreamConfig(id, maxBytes)); err != nil { return fmt.Errorf("open ingest stream: %w", err) } - q.ingest, q.maxBytes = true, maxBytes + q.ingest, q.maxBytes, q.ingestCap = true, maxBytes, maxBytes // The joins run on a budget of their own: a queue that opened but no // consumer holds fails every consumer (fail), so a slow open must not // leave them no time. @@ -371,17 +375,20 @@ func (e *EmbeddedNATS) apply(ctx context.Context, id tenant.ID, q *tenantQueue, } return nil } + prevCap := q.ingestCap if _, err := e.js.UpdateStream(resizeCtx, ingestStreamConfig(id, maxBytes)); err != nil { return fmt.Errorf("resize ingest stream: %w", err) } + q.ingestCap = maxBytes if err := e.applyDLQ(resizeCtx, id, q, maxBytes); err != nil { // The undo runs on its own budget, not the one the dead-letter call // has likely just exhausted. rollbackCtx, cancelRollback := context.WithTimeout(ctx, rollbackTimeout) defer cancelRollback() - if _, rollbackErr := e.js.UpdateStream(rollbackCtx, ingestStreamConfig(id, q.maxBytes)); rollbackErr != nil { + if _, rollbackErr := e.js.UpdateStream(rollbackCtx, ingestStreamConfig(id, prevCap)); rollbackErr != nil { return fmt.Errorf("%w (ingest stream rollback failed, so it stays at the new limit and the dlq at the previous: %w)", err, rollbackErr) } + q.ingestCap = prevCap return fmt.Errorf("%w (ingest stream restored to the previous limit)", err) } q.maxBytes = maxBytes diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index e75b62d7..3aeeb1c3 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -458,6 +458,55 @@ func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) } +// A resize whose dead-letter update fails undoes the ingest one, back to the +// cap the ingest stream had. That is not the budget applied in full: a boot +// that found the pair split applied none, and a cap of 0 would leave the +// ingest stream with no cap at all. +func TestEmbeddedNATS_SetMaxBytes_UndoRestoresTheIngestStreamsCap(t *testing.T) { + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + dir := t.TempDir() + first, err := NewEmbedded(dir) + require.NoError(t, err) + require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) + require.NoError(t, first.js.DeleteStream(ctx, "DLQ_acme")) + require.NoError(t, first.Close()) + // The dead-letter stream cannot open again: a file where its store goes. + require.NoError(t, os.WriteFile(filepath.Join(dir, "jetstream", "$G", "streams", dlqStreamName("acme")), nil, 0o600)) + + e := openEmbedded(t, dir) + require.Zero(t, e.MaxBytes("acme"), "a pair without its dead-letter stream is not at its budget") + err = e.SetMaxBytes(ctx, "acme", 16<<20) + require.ErrorContains(t, err, "ingest stream restored to the previous limit") + assert.Equal(t, int64(8<<20), streamConfig(t, e, "INGEST_acme").MaxBytes, "back at the cap it had, not unlimited") + assert.Zero(t, e.MaxBytes("acme"), "and the next call retries") +} + +// A consumer that cannot join a tenant's queue opened after it started says so +// on failed — the one report that stops the ingest worker, which would +// otherwise let the tenant's ingest answer 200 for rows nobody reads. The +// queue itself is open, so SetMaxBytes succeeds. +func TestEmbeddedNATS_Consume_ReportsAQueueItCannotJoin(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + // A durable name the client refuses: with no queue yet, nothing checks it. + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "bad.name", MaxAckPending: 10}) + require.NoError(t, err) + stop, failed, err := cons.Consume(func(*Message) {}, 4) + require.NoError(t, err) + t.Cleanup(stop) + + require.NoError(t, e.SetMaxBytes(ctx, "acme", testBudget)) + select { + case err := <-failed: + require.ErrorIs(t, err, ErrDeliveryEnded) + assert.Contains(t, err.Error(), "acme") + case <-time.After(5 * time.Second): + t.Fatal("a queue the consumer could not join was not reported") + } +} + // A budget that shrinks a tenant's dead-letter stream below what it holds // would have DiscardOld delete the oldest parked rows to fit (#532), so the // stream keeps what it holds, capped at that, and every row survives. From 07c6a91ec042fbdb33e01ea3a0379d0f418d7e54 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 20:35:20 -0400 Subject: [PATCH 006/108] fix(app): a rejected tenant keeps its replay history; review fixes --- AGENTS.md | 4 +-- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 +-- docs/src/content/docs/architecture.md | 8 ++--- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/ingest-pipeline.md | 2 +- docs/src/content/docs/sdk/admin.md | 2 +- docs/src/content/docs/sdk/streaming.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app_test.go | 29 +++++++++++++----- internal/app/wire.go | 24 ++++++++++++--- internal/ingest/sweeper.go | 8 ++--- internal/mq/mq.go | 8 ++--- internal/settings/registry.go | 18 ++++++----- internal/settings/registry_test.go | 32 ++++++++++++++++++-- internal/stream/hub.go | 10 +++--- 16 files changed, 108 insertions(+), 49 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index dbf19e84..33935174 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config @@ -45,7 +45,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`query/`** — Structured query AST types + SQL builder with schema validation, structural policy predicate/limit emission, timestamp bucketing - **`settings/`** — the settings directory, in either shape ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): flat (the four files: tenant `0` alone) or nested (one folder per tenant, never mixed). `Validate` detects the shape and checks it — `ValidateDir` per directory (strict JSON, per-file rules, cross-file role references), folder names against `tenant.Parse`, a nested finding's `File` led by its folder; `Store` is a passive holder (one tenant's adopted snapshot, typed accessors read per call); `Registry` (tenant id → `Store`) owns `Open`, the serialized `Reload`/`ReloadTenant`, the `AfterAdopt` hooks, and the fsnotify `Watch` (flat only). Flat refuses an invalid directory at boot and keeps the previous snapshot on a rejected reload; nested fails closed per tenant (a rejected folder stops being served, the rest carry on, a whole-tree reload mirrors the folders, down to none, and a finding about the root itself rejects the reload whole). Plus the embedded (`go:embed`) seed `wavehouse bootstrap` writes - **`stream/`** — SSE fan-out: rows travel POSITIONALLY, so each connection is told its projected column list in an `event: schema` frame before its first row and again on drift — **not** guaranteed after a gap-fill across a column change, which can leave a connection reading live rows against a stale list until it reconnects ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)) — (tracked per connection; replay tracks its own). The event `Hub` (registers subscribers by `(mq.Topic, role)` — one tenant's table — and evaluates each event under its own tenant's policy and schema registry; `Prune` evicts the subscribers of every tenant a reload stopped serving; `Broadcast` projects + serializes each event once per role, the #294 delivery hot path — a role carrying a row-level `filter` keeps the shared projection but delivers per subscriber, each subscriber's claims evaluated against the row, #319), `Subscriber` (per-connection outbound `Frame` queue, `Send`/`Frames`; claims fixed at construction, immutable; `Evict` asks its handler to end the stream), the `Bucket` fan-out set (`subscriberSet`, one per `(topic, role)`), the `Heartbeater` keepalive wheel, and `Metrics` (the `wavehouse_sse_*` stream instruments) -- **`tenant/`** — the tenant identifier ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): `ID` (a validated string), `Parse` (letters, digits, `_`, `-`; ≤ 64 bytes — safe as a folder name and as an MQ subject token), `Default` (`"0"`), and `Header` (`X-Tenant-ID`). Imports nothing from the rest of the repo. `api.TenantMW` resolves the header against `settings.Registry` before auth on every `/v1` route outside `/v1/ops/*` (`400` malformed, `404` unknown, a bare `503` for a nested tenant whose folder was rejected) and puts the resolved `*settings.Store` in the request context; the ops routes that address one tenant (`GET /v1/ops/pipes[/{name}]`, `POST /v1/ops/settings/reload`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh`, `POST /v1/ops/query`, `GET /v1/ops/dlq/stats`) take a strictly parsed `?tenant=` instead; handlers read it once (`api.StoreFromContext`) and pass it down as an argument, and nothing below a handler reads context. The stream hub and the ingest worker read each message's tenant off its `mq.Topic` and their getters take it; the sweeper hands the MQ each served tenant's own gap window (`gapWindows`); each served tenant has a schema registry of its own (story 6) +- **`tenant/`** — the tenant identifier ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): `ID` (a validated string), `Parse` (letters, digits, `_`, `-`; ≤ 64 bytes — safe as a folder name and as an MQ subject token), `Default` (`"0"`), and `Header` (`X-Tenant-ID`). Imports nothing from the rest of the repo. `api.TenantMW` resolves the header against `settings.Registry` before auth on every `/v1` route outside `/v1/ops/*` (`400` malformed, `404` unknown, a bare `503` for a nested tenant whose folder was rejected) and puts the resolved `*settings.Store` in the request context; the ops routes that address one tenant (`GET /v1/ops/pipes[/{name}]`, `POST /v1/ops/settings/reload`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh`, `POST /v1/ops/query`, `GET /v1/ops/dlq/stats`) take a strictly parsed `?tenant=` instead; handlers read it once (`api.StoreFromContext`) and pass it down as an argument, and nothing below a handler reads context. The stream hub and the ingest worker read each message's tenant off its `mq.Topic` and their getters take it; the sweeper hands the MQ each tenant's own gap window (`gapWindows`, a rejected tenant's included); each served tenant has a schema registry of its own (story 6) ## Key Design Decisions diff --git a/CHANGELOG.md b/CHANGELOG.md index 86dac970..90d647e0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/subscriber.go`, `internal/settings/{settings,store}.go`, `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch is shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes`, keeping no acknowledged history for a tenant no longer served, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch is shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 9b675b4c..ecc0651b 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -612,7 +612,7 @@ Opens a persistent SSE connection for real-time event streaming. Supports histor | ------ | ----------- | | `Last-Event-ID` | RFC 3339 timestamp of the last received event. If present, overrides the `since` query parameter for automatic reconnection (standard `EventSource` behavior). | -**Response:** SSE stream (`text/event-stream`). Data events include an `id:` field set to the event's `received_timestamp`. The stream opens with a `: connected` comment and emits a minimal `:` keepalive comment periodically (every 30 seconds by default), which keeps a quiet connection from being closed by a proxy; both are standard SSE comments that `EventSource` ignores (raw consumers should skip `:`-prefixed lines). When the server stops (see [Stopping](/deployment#stopping)) it ends every open stream immediately rather than holding it for the drain; `EventSource` reconnects on its own and resumes from `Last-Event-ID`. A reload that stops serving the stream's tenant — its folder removed or rejected, over a [nested settings directory](/deployment#the-nested-settings-directory) — ends that tenant's open streams the same way, and the reconnect then gets its `404` (removed) or `503` (rejected): the SDK stops on the `404` and retries the `503`, resuming from `Last-Event-ID` once the folder is back — with a hole where the tenant's history was, which the sweeper purges within a minute of the tenant no longer being served — while a browser `EventSource` treats either as fatal. A browser going cross-origin reads either refusal only when it passes CORS: it is decorated from tenant `0`'s list ([multi-tenant deployments](/deployment#multi-tenant-deployments)), so where tenant `0` is not served or its list does not admit the page's origin, the SDK sees a network error instead and keeps re-dialing. +**Response:** SSE stream (`text/event-stream`). Data events include an `id:` field set to the event's `received_timestamp`. The stream opens with a `: connected` comment and emits a minimal `:` keepalive comment periodically (every 30 seconds by default), which keeps a quiet connection from being closed by a proxy; both are standard SSE comments that `EventSource` ignores (raw consumers should skip `:`-prefixed lines). When the server stops (see [Stopping](/deployment#stopping)) it ends every open stream immediately rather than holding it for the drain; `EventSource` reconnects on its own and resumes from `Last-Event-ID`. A reload that stops serving the stream's tenant — its folder removed or rejected, over a [nested settings directory](/deployment#the-nested-settings-directory) — ends that tenant's open streams the same way, and the reconnect then gets its `404` (removed) or `503` (rejected): the SDK stops on the `404` and retries the `503`, resuming from `Last-Event-ID` once the folder is back, while a browser `EventSource` treats either as fatal. A browser going cross-origin reads either refusal only when it passes CORS: it is decorated from tenant `0`'s list ([multi-tenant deployments](/deployment#multi-tenant-deployments)), so where tenant `0` is not served or its list does not admit the page's origin, the SDK sees a network error instead and keeps re-dialing. **Row values arrive positionally, and the column names are announced separately.** Before the first row, and again whenever the column list changes, the stream sends an `event: schema` frame naming the columns of the rows that follow — in order, already reduced to what the caller's role may read. That re-announcement is **not** guaranteed after a gap-fill across a column change; see the arity note below. Every data frame's `row` array then has exactly one value per announced column, in that order. `schema` is a **named** SSE event, so a browser `EventSource` must `addEventListener('schema', …)` — it never reaches `onmessage`. A schema frame carries **no** `id:` line, so it never moves the client's `Last-Event-ID`. In the example below the table has its own `received_timestamp` **column**, which collides by name with the frame's top-level `received_timestamp` **field** — they are different values: the field is when WaveHouse received the event, the row slot is that column as published (`null` where the record omitted it, which ClickHouse replaces with the column's default on insert). @@ -745,7 +745,7 @@ Triggers an immediate re-discovery of the `?tenant=`'s ClickHouse table schemas #### `GET /v1/ops/dlq/stats` — DLQ Statistics -Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, for as long as its queue is kept. The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream exists from the moment the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. +Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, since its queue is kept (nothing deletes it). The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream exists from the moment the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. **Error responses:** diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 549622f6..30a04ae2 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -139,7 +139,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy. On a bulk-insert failure the batch is re-inserted row by row — except a batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling), which no row could pass and `parkBatch` takes to the DLQ switch whole, logging once per batch rather than twice per row; rows that succeed are acked, and only the rows that fail again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. - **types.go** — `EventMessage` struct (TableName, Scope — reserved, always empty today, ReceivedTimestamp, Format, Columns, Row; `Format` is `FormatJSONCompactEachRow` and `Row` is one positional line whose slots `Columns` names) and `BufferConsumerName` constant, shared across API handlers and the ingest pipeline. - **compact.go** — `EncodeCompactRow`, the positional row encoder every published row goes through, rendering one record over the table's **insertable** columns in declaration order. Serialization only: it validates nothing and judges no value. -- **sweeper.go** — `Sweeper` implements the Active Sweeper pattern. It runs every minute and asks the MQ (`mq.Purger.PurgeAcked`) to drop the ingest events that are **both** ACKed by the buffer consumer (written to ClickHouse) **and** older than the gap window (re-read every sweep: each served tenant's own `stream.gap_window_minutes` — `internal/app`'s `gapWindows` — and none for a tenant no longer served). Finding the purge point is `internal/mq`'s (`purge.go`). +- **sweeper.go** — `Sweeper` implements the Active Sweeper pattern. It runs every minute and asks the MQ (`mq.Purger.PurgeAcked`) to drop the ingest events that are **both** ACKed by the buffer consumer (written to ClickHouse) **and** older than the gap window (re-read every sweep: each tenant's own `stream.gap_window_minutes`, a rejected tenant's as its folder last had it (unbounded for one rejected since boot) — `internal/app`'s `gapWindows` — and none for a removed tenant). Finding the purge point is `internal/mq`'s (`purge.go`). ### `mq/` — Message Queue @@ -190,7 +190,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `tenant/` — Tenant Identifier -- **tenant.go** — `ID`, a validated string (never a number: a 19-digit id already rounds as a float64), and `Parse`, the one grammar that makes an id safe both as a folder name and as a message-queue subject token: ASCII letters, digits, `_`, `-`, at most `MaxLen` (64) bytes. `Default` (`"0"`) is the tenant a request without the header resolves to; `Header` is `X-Tenant-ID`. The package imports nothing from the rest of the repository, so any package can name a tenant. HTTP handlers receive the tenant as its resolved `*settings.Store`, which knows its id (`Store.Tenant`) for the topics they publish and subscribe on; the stream hub and the ingest worker read each event's tenant off its `mq.Topic` — the leading subject token — and their settings getters take it as a parameter, which `internal/app` resolves through the registry; the sweeper hands the MQ each served tenant's own gap window (`gapWindows`); each served tenant has a schema registry of its own, built with its id (story 6). +- **tenant.go** — `ID`, a validated string (never a number: a 19-digit id already rounds as a float64), and `Parse`, the one grammar that makes an id safe both as a folder name and as a message-queue subject token: ASCII letters, digits, `_`, `-`, at most `MaxLen` (64) bytes. `Default` (`"0"`) is the tenant a request without the header resolves to; `Header` is `X-Tenant-ID`. The package imports nothing from the rest of the repository, so any package can name a tenant. HTTP handlers receive the tenant as its resolved `*settings.Store`, which knows its id (`Store.Tenant`) for the topics they publish and subscribe on; the stream hub and the ingest worker read each event's tenant off its `mq.Topic` — the leading subject token — and their settings getters take it as a parameter, which `internal/app` resolves through the registry; the sweeper hands the MQ each tenant's own gap window (`gapWindows`, a rejected tenant's included); each served tenant has a schema registry of its own, built with its id (story 6). ### `chconn/` — ClickHouse Connection Pools @@ -255,7 +255,7 @@ Ingest worker pipeline (StartIngestWorker): Active Sweeper (async goroutine, every 60s), on each tenant's stream: → Read buffer consumer's AckFloor (highest contiguous ACKed seq) → Binary search for first message within that tenant's own gap window - (none for a tenant no longer served) + (a rejected tenant's as its folder last had it, unbounded for one rejected since boot; none for a removed tenant) → Purge target = MIN(ack_floor + 1, gap_window_seq) → Purge all messages below target from JetStream ``` diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 2029bc0d..6ad7f34e 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -383,7 +383,7 @@ That is the layout a control plane writes. Each folder's `clickhouse` block is i The folder name is the tenant id, and each folder is a complete settings directory: everything on the [Settings Directory](/settings-directory) page applies to it as written, except where the rules below say otherwise. The two shapes don't mix — a folder beside the four files, or a loose file beside the folders, is a validation error — and a running server keeps the shape it booted with, so switching is stop, restructure, start. The dedupe store needs no restructuring: it keys every tenant's seen ids by tenant, and the four files are tenant `0`, as a `0` folder is. Dot-prefixed entries are ignored in either shape. `wavehouse validate` checks either shape with the same exit codes; a finding in a nested directory names its folder (`acme/policies.json`), and a folder whose name is not a tenant id is a finding of its own — that folder is skipped, and the rest of the directory still loads. -**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). Its message queue is kept, at the budget it last had, but the history that gap-fill replays is purged from it at the next sweep, as a removed tenant's is, so a stream resumed after the fix has a hole where that history was. The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. +**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). Its message queue is kept, at the budget it last had, and so is the history that gap-fill replays, for the `stream.gap_window_minutes` its folder last had (all of it, for a folder rejected since the server started, whose window the server never read): a stream resumed after a fix within that window picks up where it left off. The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. **Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted — its message queue is kept at the budget it last had, and only the history that gap-fill replays goes from it, at the next sweep — so restoring the folder restores the tenant, seen ids and parked rows included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index 72447024..448fa6a0 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -228,7 +228,7 @@ Several layers throttle the pipeline, inner to outer: ## The Active Sweeper -The worker advances the consumer's `AckFloor` by acking; the sweep observes it to decide what is safe to purge. They never call each other — the consumer's `AckFloor` is their only contract. The sweeper (`internal/ingest`) owns the schedule and the window: each tick it calls `mq.Purger.PurgeAcked(buffer-consumer, cutoffs)` with each served tenant's cutoff at now − its own `stream.gap_window_minutes`; a tenant no longer served — its folder removed or rejected — is given none, and keeps none of the history it has acknowledged. The steps after the tick below are the embedded broker's implementation of that call, run on each tenant's stream at that tenant's cutoff. +The worker advances the consumer's `AckFloor` by acking; the sweep observes it to decide what is safe to purge. They never call each other — the consumer's `AckFloor` is their only contract. The sweeper (`internal/ingest`) owns the schedule and the window: each tick it calls `mq.Purger.PurgeAcked(buffer-consumer, cutoffs)` with each tenant's cutoff at now − its own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it, or one before anything it holds if its folder has been rejected since boot, so its clients resume once the folder is fixed; a removed tenant is given none, and keeps none of the history it has acknowledged. The steps after the tick below are the embedded broker's implementation of that call, run on each tenant's stream at that tenant's cutoff. ```mermaid flowchart TD diff --git a/docs/src/content/docs/sdk/admin.md b/docs/src/content/docs/sdk/admin.md index 59db0e5a..1dee059d 100644 --- a/docs/src/content/docs/sdk/admin.md +++ b/docs/src/content/docs/sdk/admin.md @@ -72,7 +72,7 @@ const { data } = await wh.dlq.list({ tenant: 'acme' }); const { data: clicks } = await wh.dlq.table('clicks', { tenant: 'acme' }); ``` -`wh.dlq.stream()` exists in the API but is **not yet functional**: there is no server-side DLQ stream today (the SSE bridge only carries `ingest.>` subjects), so it connects and receives no events rather than failing. Live DLQ streaming is tracked in [#197](https://github.com/Wave-RF/WaveHouse/issues/197). +`wh.dlq.stream()` exists in the API but is **not yet functional**: there is no server-side SSE route for dead-lettered events today (the SSE bridge only carries `ingest.>` subjects), so it connects and receives no events rather than failing. Live DLQ streaming is tracked in [#197](https://github.com/Wave-RF/WaveHouse/issues/197). --- diff --git a/docs/src/content/docs/sdk/streaming.md b/docs/src/content/docs/sdk/streaming.md index 2b37667a..94a7e334 100644 --- a/docs/src/content/docs/sdk/streaming.md +++ b/docs/src/content/docs/sdk/streaming.md @@ -116,7 +116,7 @@ A dropped stream reconnects on a jittered exponential backoff, capped at 30s, an :::caution[Resumption is at-least-once, and time-bounded] Delivery across a reconnect is **at-least-once**. The `Last-Event-ID` the client sends is the last event's `received_timestamp`, and the server replays from that instant *inclusively* — so the last event you already saw, and anything sharing its timestamp, arrives again. The SDK does not deduplicate live frames — `liveQuery()` makes one pass at the backfill seam, and only under an ascending order ([#449](https://github.com/Wave-RF/WaveHouse/issues/449)) — so key on `timestamp` plus your own row identity if duplicates matter. -Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. So does a stream a [nested server](/deployment#the-nested-settings-directory) ended because its tenant's folder was rejected, once the folder is fixed: the sweeper purges a tenant's history within a minute of the tenant no longer being served. The same silence applies across a server upgrade to this release: the server deletes the previous release's queue at boot, so a replay spanning the upgrade omits the events published before it, without an error — backfill over REST if you need them. +Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. The same silence applies across a server upgrade to this release: the server deletes the previous release's queue at boot, so a replay spanning the upgrade omits the events published before it, without an error — backfill over REST if you need them. **A column-set change across a gap-fill is a known limitation.** If the table's columns change while you are connected *and* your client replays across that change, live rows arriving after the replay may not be preceded by a fresh `event: schema` frame until the columns next change or you reconnect. The SDK drops a row whose **length** disagrees with the list it was last told, rather than zipping it under the wrong names — so an added or removed column costs you rows, not wrong ones. A **same-length** change is the residual case the arity check cannot see: a `RENAME COLUMN`, or a drop paired with an add, zips values under the wrong names until the next announcement. Reconnecting resynchronizes either way. Full schema-change handling is deferred to the schema-versioning work ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)). ::: diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index a0808514..99b1f679 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -221,7 +221,7 @@ A tenant's dead-letter stream exists from the moment the tenant is first served ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume. A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume — counting every tenant ever served on it, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream (up to a tenth of its last budget). A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. ## Streaming diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 8ef1f7c5..17e92812 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -730,10 +730,12 @@ func gapWindow(minutes int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": minutes}} } -// Each tenant being served keeps its own stream.gap_window_minutes, since -// each has a queue of its own; a rejected tenant is not served, so it is not -// named and keeps no history (mq.Purger.PurgeAcked). A flat directory's single -// tenant gets exactly its own window. +// Each tenant keeps its own stream.gap_window_minutes, since each has a queue +// of its own — a rejected tenant the window its folder last had, so its +// clients resume once the folder is fixed, and everything while that window +// is unknown. A removed tenant is not named and keeps no history +// (mq.Purger.PurgeAcked). A flat directory's single tenant gets exactly its +// own window. func TestGapWindows(t *testing.T) { open := func(t *testing.T, dir string) *settings.Registry { t.Helper() @@ -754,11 +756,24 @@ func TestGapWindows(t *testing.T) { rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) tenants.Reload("test") - assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants)) + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "globex": 60 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants), + "a rejected tenant keeps the window its folder last had") + + require.NoError(t, os.RemoveAll(filepath.Join(root, "globex"))) + tenants.Reload("test") + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants), + "a removed tenant keeps none") }) - t.Run("no tenant served names none", func(t *testing.T) { - assert.Empty(t, gapWindows(open(t, writeNestedSettings(t, map[string]map[string]any{"acme": invalidQuery})))) + t.Run("a folder rejected since boot keeps everything", func(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": invalidQuery}) + tenants := open(t, root) + assert.Equal(t, map[tenant.ID]time.Duration{"acme": keepEverything}, gapWindows(tenants)) + + rewriteSettings(t, filepath.Join(root, "acme"), gapWindow(15)) + tenants.Reload("test") + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute}, gapWindows(tenants), + "its own window once its folder validates") }) } diff --git a/internal/app/wire.go b/internal/app/wire.go index e08f4a0e..ed3b98ec 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -6,6 +6,7 @@ import ( "fmt" "log/slog" "maps" + "math" "net" "net/http" "os" @@ -130,18 +131,31 @@ func shortestKeepalive(tenants *settings.Registry) (period time.Duration, bucket return period, buckets } -// gapWindows is the history the sweeper keeps for each tenant being served: -// its own stream.gap_window_minutes, since each tenant's events have a queue -// of their own. A tenant it does not name — removed or rejected — keeps no -// history (mq.Purger.PurgeAcked). +// gapWindows is the history the sweeper keeps for each tenant: its own +// stream.gap_window_minutes, since each tenant's events have a queue of their +// own — for a rejected tenant, the window its folder last had, because a +// rejection is the common reload failure (a typo, fixed minutes later) and +// its clients resume from Last-Event-ID once it is served again. A removed +// tenant is not named, so it keeps no history (mq.Purger.PurgeAcked). func gapWindows(tenants *settings.Registry) map[tenant.ID]time.Duration { windows := map[tenant.ID]time.Duration{} - for id, store := range tenants.All() { + for id, store := range tenants.Known() { + if store == nil { + windows[id] = keepEverything + continue + } windows[id] = store.GapWindow() } return windows } +// keepEverything is the window of a tenant whose folder has been rejected +// since boot: this process has never read its stream.gap_window_minutes, so +// none of the history its queue holds is known to be past it. A rejected +// tenant is sent no new events, so what it keeps is what its queue held at +// boot. +const keepEverything = time.Duration(math.MaxInt64) + // served reports whether the registry is serving tenant id: what the // per-tenant resources — verifiers, dedupe stores, open streams — are pruned // by once a reload removes or rejects their tenant. diff --git a/internal/ingest/sweeper.go b/internal/ingest/sweeper.go index 367e22af..b0d4a2d1 100644 --- a/internal/ingest/sweeper.go +++ b/internal/ingest/sweeper.go @@ -22,10 +22,10 @@ import ( // (see mq.Purger). type Sweeper struct { purger mq.Purger - // gapWindows is the history to keep for each tenant being served, read on - // every sweep so a reload of stream.gap_window_minutes applies from the - // next sweep without a restart. A tenant it does not name — one removed - // or rejected — keeps no history (mq.Purger.PurgeAcked). + // gapWindows is the history to keep for each tenant, read on every sweep + // so a reload of stream.gap_window_minutes applies from the next sweep + // without a restart. A tenant it does not name keeps no history + // (mq.Purger.PurgeAcked). gapWindows func() map[tenant.ID]time.Duration } diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 663bde3d..47897978 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -276,10 +276,10 @@ type Purger interface { // its first unacked event) AND stored before that tenant's cutoff in // olderThan. Either bound alone keeps the event: unacked events are not // yet written, and recent ones are still needed for replay. A tenant - // olderThan does not name — one no longer served — keeps no history: - // everything it has acknowledged goes. Reports whether anything was - // removed. ErrConsumerNotFound when the consumer has not been created on - // some tenant's queue; the other tenants' are purged all the same. + // olderThan does not name keeps no history: everything it has + // acknowledged goes. Reports whether anything was removed. + // ErrConsumerNotFound when the consumer has not been created on some + // tenant's queue; the other tenants' are purged all the same. PurgeAcked(ctx context.Context, consumer string, olderThan map[tenant.ID]time.Time) (purged bool, err error) } diff --git a/internal/settings/registry.go b/internal/settings/registry.go index 663e0a23..e624a780 100644 --- a/internal/settings/registry.go +++ b/internal/settings/registry.go @@ -154,13 +154,17 @@ func (r *Registry) All() iter.Seq2[tenant.ID, *Store] { } // Known iterates over every tenant the registry holds, served or rejected, -// in id order — for a consumer that must keep a rejected tenant's resources -// current too: the tenant comes back into service with them, and a rejection -// is the common reload failure (a typo, fixed and reloaded minutes later). -func (r *Registry) Known() iter.Seq[tenant.ID] { - return func(yield func(tenant.ID) bool) { - for _, id := range slices.Sorted(maps.Keys(*r.tenants.Load())) { - if !yield(id) { +// in id order, each with the store holding its last adopted settings — nil +// for a tenant whose folder has not validated since boot — for a consumer +// that must keep a rejected tenant's resources current too: the tenant comes +// back into service with them, and a rejection is the common reload failure +// (a typo, fixed and reloaded minutes later). A rejected tenant's store +// serves no request; it is handed out for what those resources read of it. +func (r *Registry) Known() iter.Seq2[tenant.ID, *Store] { + return func(yield func(tenant.ID, *Store) bool) { + tenants := *r.tenants.Load() + for _, id := range slices.Sorted(maps.Keys(tenants)) { + if !yield(id, tenants[id].store) { return } } diff --git a/internal/settings/registry_test.go b/internal/settings/registry_test.go index 0bdff569..06e2a3b9 100644 --- a/internal/settings/registry_test.go +++ b/internal/settings/registry_test.go @@ -5,7 +5,6 @@ import ( "log/slog" "os" "path/filepath" - "slices" "testing" "github.com/stretchr/testify/assert" @@ -307,18 +306,45 @@ func TestRegistry_HooksRunOnAReloadThatAdoptsNothing(t *testing.T) { } // Known is every tenant the registry holds, rejected ones included, in id -// order: what a resource a rejected tenant comes back to is kept current for. +// order, each with the store of its last adopted settings: what a resource a +// rejected tenant comes back to is kept current from. func TestRegistry_Known(t *testing.T) { t.Parallel() root := writeTree(t, map[string]map[string]string{"globex": maxRowsFiles(222), "acme": maxRowsFiles(111), "broken": brokenFiles()}) reg, _ := Open(root) require.NotNil(t, reg) - assert.Equal(t, []tenant.ID{"acme", "broken", "globex"}, slices.Collect(reg.Known())) + known := func() ([]tenant.ID, map[tenant.ID]*Store) { + var ids []tenant.ID + stores := map[tenant.ID]*Store{} + for id, store := range reg.Known() { + ids = append(ids, id) + stores[id] = store + } + return ids, stores + } + ids, stores := known() + assert.Equal(t, []tenant.ID{"acme", "broken", "globex"}, ids) + assert.Nil(t, stores["broken"], "a folder that has not validated since boot has no settings to hand out") + acme, _ := reg.For("acme") + assert.Same(t, acme, stores["acme"]) var served []tenant.ID for id := range reg.All() { served = append(served, id) } assert.Equal(t, []tenant.ID{"acme", "globex"}, served, "All leaves the rejected tenant out; Known does not") + + // A tenant rejected after an adoption still comes with that adoption's + // settings: the store keeps its last document. + for name, content := range brokenFiles() { + require.NoError(t, os.WriteFile(filepath.Join(root, "acme", name), []byte(content), 0o600)) + } + reg.Reload("test") + _, ok := reg.For("acme") + require.False(t, ok) + _, stores = known() + require.Same(t, acme, stores["acme"]) + assert.Equal(t, 111, stores["acme"].DefaultMaxRows()) + // Stopping early is the iterator's contract, not the caller's problem. for id := range reg.Known() { assert.Equal(t, tenant.ID("acme"), id) diff --git a/internal/stream/hub.go b/internal/stream/hub.go index e7ce2d4d..b13221bf 100644 --- a/internal/stream/hub.go +++ b/internal/stream/hub.go @@ -43,8 +43,9 @@ type Hub struct { // the default implementation, which delegates to ResolvedPermissions.RowVisible // — today's behavior unchanged. Wired once before the Hub serves traffic and // not safe to mutate afterwards: rowAdmitted reads it from the consumer - // goroutine and from SSE handler goroutines without holding h.mu. Every - // delivery path reaches it through rowAdmitted, never directly. + // goroutines (one per tenant) and from SSE handler goroutines without + // holding h.mu. Every delivery path reaches it through rowAdmitted, never + // directly. RowEvaluator RowEvaluator } @@ -393,9 +394,8 @@ func newEventView(raw []byte) *eventView { // The hub is a second consumer of the same events as the ingest worker, and // acks independently of it, so a format only the worker refuses would stream // to clients while the worker parks it on the DLQ. Refusing it here keeps the - // two readers agreeing on what the bytes mean. Today only a pre-v2 envelope - // declares anything else, and it would fail pairing anyway on its empty - // column list — this is what holds once a second format exists. + // two readers agreeing on what the bytes mean. Today every envelope + // declares that one format — this is what holds once a second exists. if ev.evt.Format != ingest.FormatJSONCompactEachRow { return ev } From 04aad9ffb4751b8719637b1bb214cb683154fd8a Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 21:02:27 -0400 Subject: [PATCH 007/108] fix(mq): share the hub bridge's fetch-ahead across tenants; review fixes --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 ++-- docs/src/content/docs/architecture.md | 4 ++-- docs/src/content/docs/settings-directory.mdx | 2 +- internal/mq/embedded.go | 10 ++++++---- internal/mq/embedded_test.go | 16 ++++++++++++++++ internal/mq/mq.go | 4 +++- internal/settings/settings.go | 4 ++-- 8 files changed, 33 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 90d647e0..76564ceb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch is shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index ecc0651b..df65cfe4 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -745,7 +745,7 @@ Triggers an immediate re-discovery of the `?tenant=`'s ClickHouse table schemas #### `GET /v1/ops/dlq/stats` — DLQ Statistics -Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, since its queue is kept (nothing deletes it). The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream exists from the moment the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. +Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, since its queue is kept (nothing deletes it). The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream is opened when the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. **Error responses:** @@ -754,7 +754,7 @@ Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant] | 401 | `{"error":"invalid token"}` / `{"error":"token expired"}` | A present-but-invalid/expired token was supplied and denied (the gate surfaces the token reason) | | 400 | `{"error":"invalid query string: …"}` / `{"error":"invalid ?tenant: …"}` | The query string does not parse (`?tenant=acme;x=1`, a bad `%` escape), or `tenant` is empty, repeated, or not a tenant id | | 403 | `{"error":"forbidden"}` | Caller's role is not the policy `admin_role` (`"admin"` by default) | -| 404 | `{"error":"no dead-letter queue for tenant: "}` | The tenant has no dead-letter queue: it has never been served on this data directory, or the id names no tenant | +| 404 | `{"error":"no dead-letter queue for tenant: "}` | The tenant has no dead-letter queue: it has never been served on this data directory, its queue could not be opened (see [Message Queue](/settings-directory#message-queue)), or the id names no tenant | | 500 | `{"error":"stream info failed"}` | NATS JetStream stream-info lookup failed | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while tenant `0`'s JWKS has not been fetched yet (the ops tree verifies as tenant `0`); refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 30a04ae2..51ab00d0 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -145,7 +145,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. -- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. +- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. It runs on each tenant's stream at that tenant's cutoff. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. - **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, or is refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 99b1f679..caf5b685 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -217,7 +217,7 @@ A failed batch insert is retried row by row; a row that fails again on its own i For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in its ingest queue, redelivered for as long as the tenant is away and stopping the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging that queue. -A tenant's dead-letter stream exists from the moment the tenant is first served (an empty stream costs nothing) and the stats endpoint is always registered — the switch is purely behavioral, which is what makes it safe to reload. +A tenant's dead-letter stream is opened when the tenant is first served (an empty stream costs nothing) and the stats endpoint is always registered — the switch is purely behavioral, which is what makes it safe to reload. ## Message Queue diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index c3762a8f..2a990960 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -547,9 +547,11 @@ func wrapMsg(ctx context.Context, m jetstream.Msg) *Message { // Subscribe holds a durable explicit-ack consumer named consumerName on every // tenant's queue, those opened later included, and delivers each message to -// handler with the trace context its headers carry, until ctx is done. A -// tenant's queue that cannot be joined when it opens is logged: its events -// reach handler from the next boot. +// handler with the trace context its headers carry, until ctx is done. It +// fetches the client's default number of messages ahead across the tenants +// together (see fanIn.share), so what sits client-side does not grow with +// the tenants. A tenant's queue that cannot be joined when it opens is +// logged: its events reach handler from the next boot. func (e *EmbeddedNATS) Subscribe(ctx context.Context, consumerName string, handler func(msg *Message) error) error { f := e.newFanIn(ctx, jetstream.ConsumerConfig{Durable: consumerName, AckPolicy: jetstream.AckExplicitPolicy}) f.fail = func(err error) { @@ -563,7 +565,7 @@ func (e *EmbeddedNATS) Subscribe(ctx context.Context, consumerName string, handl if err := handler(msg); err != nil { _ = msg.Nak() } - }, 0, false) + }, jetstream.DefaultMaxMessages, false) if err != nil { return fmt.Errorf("consume: %w", err) } diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 3aeeb1c3..f22fb7b0 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -1070,6 +1070,22 @@ func TestFanIn_SharesThePrefetch(t *testing.T) { assert.Zero(t, (&fanIn{handles: handles(3)}).share(), "0 leaves the client default") } +// The hub bridge's fetch-ahead is the client default split across the +// tenants' queues, like the worker's prefetch, so what it holds client-side +// does not grow with the number of tenants. +func TestEmbeddedNATS_Subscribe_SharesTheClientDefault(t *testing.T) { + e := newTestEmbedded(t, "acme", "globex") + ctx, cancel := context.WithCancel(t.Context()) + defer cancel() + require.NoError(t, e.Subscribe(ctx, "hub-bridge", func(*Message) error { return nil })) + + e.mu.Lock() + defer e.mu.Unlock() + require.Len(t, e.consumers, 1) + assert.Equal(t, jetstream.DefaultMaxMessages, e.consumers[0].prefetch) + assert.Equal(t, jetstream.DefaultMaxMessages/2, e.consumers[0].share()) +} + // Nothing lands on the default tenant by omission (#583): the tenant is a // required token, checked against its grammar before anything is sent. func TestEmbeddedNATS_Publish_RefusesATopicWithoutATenant(t *testing.T) { diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 47897978..6626cfa8 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -164,7 +164,9 @@ type Subscriber interface { // consumer named consumerName, held on every tenant's queue — those // opened after Subscribe included. The handler runs on one delivery // goroutine per tenant, one message at a time, so it must be safe to - // call concurrently for different tenants. + // call concurrently for different tenants. The messages fetched ahead of + // it are a fixed number split across the tenants, as Consumer.Consume's + // prefetch is, so they do not grow with the number of tenants. // // CONTRACT: If the handler intends to return an error to trigger automatic // redelivery, it MUST NOT manually call msg.Ack() or msg.Nak() beforehand. diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 2ce9b120..7da6c3fa 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -165,8 +165,8 @@ type TableDedupe struct { // DLQConfig gates the Dead Letter Queue: whether a row that still fails // after the row-by-row isolation retry is parked on the tenant's dead-letter // queue (and its original acked) or left unacked to be redelivered -// indefinitely. The queue exists from the moment the tenant is first served — -// empty until something lands on it — so the switch is purely behavioral and +// indefinitely. The queue is opened when the tenant is first served — empty +// until something lands on it — so the switch is purely behavioral and // resolves per table through the same override cascade as dedupe. type DLQConfig struct { Enabled *bool `json:"enabled"` From 57870c4d19dea281845e85c386ee580894070645 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 21:33:51 -0400 Subject: [PATCH 008/108] fix(mq): publish only into a queue the broker recorded open; review fixes --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 +- docs/src/content/docs/settings-directory.mdx | 4 +- internal/api/ingest.go | 4 +- internal/app/app.go | 10 ++-- internal/app/wire.go | 48 +++++------------- internal/mq/embedded.go | 51 +++++++++++++++----- internal/mq/embedded_test.go | 42 ++++++++++++++++ 9 files changed, 105 insertions(+), 62 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 33935174..3d780d92 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config diff --git a/CHANGELOG.md b/CHANGELOG.md index 76564ceb..c24d0df9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 51ab00d0..9da998d1 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -148,7 +148,7 @@ The **only** package that imports NATS/JetStream — a `depguard` rule in `.gola - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. It runs on each tenant's stream at that tenant's cutoff. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, or is refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index caf5b685..d34c0a65 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -221,7 +221,9 @@ A tenant's dead-letter stream is opened when the tenant is first served (an empt ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume — counting every tenant ever served on it, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream (up to a tenth of its last budget). A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. + +**Sizing the volume.** Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume — counting every tenant ever served on it, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream (up to a tenth of its last budget). A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. ## Streaming diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 029799b6..10daddea 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -707,10 +707,10 @@ func (h *IngestHandler) processRecord( slog.DebugContext(ctx, "publishing event to the ingest queue", "table", table, "scope", scope) if err := h.Publisher.Publish(ctx, mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope}, payload); err != nil { if errors.Is(err, mq.ErrQueueFull) { - slog.WarnContext(ctx, "ingest queue is full", "error", err, "table", table, "scope", scope) + slog.WarnContext(ctx, "ingest queue is full", "tenant", store.Tenant(), "error", err, "table", table, "scope", scope) return false, nil, &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} } - slog.ErrorContext(ctx, "failed to publish to the ingest queue", "error", err, "table", table, "scope", scope) + slog.ErrorContext(ctx, "failed to publish to the ingest queue", "tenant", store.Tenant(), "error", err, "table", table, "scope", scope) return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "publish failed"} } diff --git a/internal/app/app.go b/internal/app/app.go index 2131942c..a7aec9d1 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -18,7 +18,7 @@ // handed whole to each component's wiring function, which derives the // per-call getters the internal packages take: keyed by the request's store // for the handlers, by tenant id for the async paths (perTenant), and fixed -// to the default tenant for the ops gate of a flat directory (defaultSetting). +// to the default tenant for the ops gate of a flat directory (defaultPolicy). package app import ( @@ -30,7 +30,6 @@ import ( "net/http" "os" "os/signal" - "sync/atomic" "time" "golang.org/x/sync/errgroup" @@ -89,11 +88,8 @@ type App struct { listener net.Listener // tenants is the registry every tenant-aware path resolves through, and - // the owner of every reload. defaultStore is tenant 0's store as of its - // last adoption, which the ops gate of a flat directory reads its admin - // role from (defaultSetting). - tenants *settings.Registry - defaultStore atomic.Pointer[settings.Store] + // the owner of every reload. + tenants *settings.Registry // policies is the default tenant's policy, for the ops gate of a flat // directory. policies policy.Source diff --git a/internal/app/wire.go b/internal/app/wire.go index ed3b98ec..968d5f34 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -66,50 +66,24 @@ func (a *App) wireSettings() error { return fmt.Errorf("settings directory %s invalid, refusing to start — findings above; `wavehouse validate` reproduces them, `wavehouse bootstrap` writes a starter directory", a.cfg.Settings.Dir) } a.tenants = tenants - // Registered first: hooks run in registration order, so every reload - // updates the tracked store before any other hook runs. - a.trackDefaultStore() - a.onDefaultAdopt(a.trackDefaultStore) - a.policies = func() *policy.Policy { return defaultSetting(a, (*settings.Store).Policy) } + a.policies = func() *policy.Policy { return defaultPolicy(tenants) } if !tenants.Nested() && a.policies() == nil { slog.Warn("no policy adopted — every token-based request is denied until policies.json defines one (fail closed)") } return nil } -// trackDefaultStore remembers tenant 0's store as of its last adoption. The -// registry stops handing out a rejected tenant's store and forgets a removed -// one, but the store keeps its last adopted document either way — and that is -// what defaultSetting goes on reading. -func (a *App) trackDefaultStore() { - if store, ok := a.tenants.For(tenant.Default); ok { - a.defaultStore.Store(store) +// defaultPolicy is the default tenant's access-control policy, which the ops +// gate of a flat directory reads its admin role from per request. There +// tenant 0 is the whole directory, always served: a reload that fails keeps +// the previous document. A nested directory's ops gate reads no policy at all +// (api.NewRouter). +func defaultPolicy(tenants *settings.Registry) *policy.Policy { + store, ok := tenants.For(tenant.Default) + if !ok { + return nil } -} - -// defaultSetting reads one setting of the default tenant: the admin role the -// ops gate of a flat directory reads per request. It reads tenant 0's last -// adopted document, so a 0 folder a reload rejected or removed leaves its -// reader as it was. A nested directory that has never served a tenant 0 reads -// T's zero value, and its ops gate reads no policy at all. -func defaultSetting[T any](a *App, get func(*settings.Store) T) T { - store := a.defaultStore.Load() - if store == nil { - var zero T - return zero - } - return get(store) -} - -// onDefaultAdopt registers fn to run after each reload that adopts the -// default tenant, so a nested directory's other tenants never move what -// follows it, and a rejected 0 folder leaves that as it was. -func (a *App) onDefaultAdopt(fn func()) { - a.tenants.AfterAdopt(func(adopted []tenant.ID) { - if slices.Contains(adopted, tenant.Default) { - fn() - } - }) + return store.Policy() } // shortestKeepalive is the shape of the one keepalive wheel every tenant's diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 2a990960..bebcfa63 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -57,15 +57,20 @@ type EmbeddedNATS struct { conn *nats.Conn js jetstream.JetStream - // mu guards queues and consumers, and serializes opening or resizing a - // tenant's queue with registering a consumer, so a queue opened while a - // consumer registers is never missed by it. It is held across the - // JetStream calls that open or resize a queue. + // mu guards queues, consumers and writes to opened, and serializes + // opening or resizing a tenant's queue with registering a consumer, so a + // queue opened while a consumer registers is never missed by it. It is + // held across the JetStream calls that open or resize a queue. mu sync.Mutex queues map[tenant.ID]*tenantQueue // consumers are the durable consumers held on every tenant's queue, each // joined to a queue as it opens. consumers []*fanIn + // opened holds the tenants whose queue has both streams and every + // registered consumer joined — what Publish trusts, rather than a stream + // answering: an open that gave up can leave behind a stream JetStream goes + // on to create, which no consumer holds. Written under mu, read without it. + opened sync.Map // tenant.ID → struct{} } // tenantQueue is what the broker knows of one tenant's queue. @@ -219,6 +224,7 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { if q.ingest && ok && (d.limit == tenth || guarded) { q.maxBytes = q.asked } + e.record(id, q) } return nil } @@ -267,6 +273,19 @@ func (e *EmbeddedNATS) ingestTenants() []tenant.ID { return ids } +// record brings opened in line with what the broker knows of tenant id's +// queue. Both streams known means every consumer holds the queue too: apply +// joins the consumers to a queue it opens before this records it, and a +// consumer registered later joins every ingest stream there is. Under e.mu +// (or before e is shared). +func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { + if q.ingest && q.dlq { + e.opened.Store(id, struct{}{}) + } else { + e.opened.Delete(id) + } +} + // ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard // append-only log; the Active Sweeper handles message purging. MaxBytes caps // the tenant's share of the disk. DiscardNew rejects new messages when full, @@ -353,6 +372,7 @@ func (e *EmbeddedNATS) SetMaxBytes(ctx context.Context, id tenant.ID, maxBytes i // apply brings tenant id's queue to maxBytes: opening it when its ingest // stream is missing, resizing it otherwise (see SetMaxBytes). Under e.mu. func (e *EmbeddedNATS) apply(ctx context.Context, id tenant.ID, q *tenantQueue, maxBytes int64) error { + defer e.record(id, q) resizeCtx, cancel := context.WithTimeout(ctx, resizeTimeout) defer cancel() if !q.ingest { @@ -426,9 +446,10 @@ func (e *EmbeddedNATS) applyDLQ(ctx context.Context, id tenant.ID, q *tenantQueu } // reopen opens tenant id's queue at the budget last asked for it, for a -// publish or park that found one of its streams missing. errNoQueue when no -// budget has been asked for the tenant yet: a reload can make a tenant -// resolvable an instant before its budget arrives. +// publish that finds the queue not recorded open, or a publish or park that +// found one of its streams missing. errNoQueue when no budget has been asked +// for the tenant yet: a reload can make a tenant resolvable an instant before +// its budget arrives. // // It runs detached from ctx's cancellation, bounded by its own timeouts: // ctx is one caller's — an ingest request — while the queue is every @@ -442,6 +463,7 @@ func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { if q == nil || q.asked == 0 { return fmt.Errorf("tenant %s: %w", id, errNoQueue) } + defer e.record(id, q) // What is missing is asked of JetStream rather than read off the flags, // which may still say the stream the publish just missed exists — or it // may be back already, opened by a caller that held mu first. @@ -467,15 +489,22 @@ func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { // Publish stores data on topic's ingest subject, in its tenant's queue. A // topic without a valid tenant is refused before anything is sent (see // subject). A tenant with no queue has one opened at the budget last asked -// for it (see SetMaxBytes). A queue that cannot be opened — none asked for -// yet, or JetStream refused it — and a queue at its byte budget (DiscardNew) -// are reported as ErrQueueFull: either way the tenant's queue takes nothing -// now, and a retry is the caller's answer. +// for it (see SetMaxBytes) — and so does one whose stream exists but whose +// queue the broker has not recorded open, since no consumer may hold that +// stream. A queue that cannot be opened — none asked for yet, or JetStream +// refused it — and a queue at its byte budget (DiscardNew) are reported as +// ErrQueueFull: either way the tenant's queue takes nothing now, and a retry +// is the caller's answer. func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, opts ...PublishOpt) error { subj, err := subject(ingestPrefix, topic) if err != nil { return err } + if _, ok := e.opened.Load(topic.Tenant); !ok { + if openErr := e.reopen(ctx, topic.Tenant); openErr != nil { + return fmt.Errorf("%w: %w", ErrQueueFull, openErr) + } + } err = e.publish(ctx, subj, data, opts) if errors.Is(err, jetstream.ErrNoStreamResponse) { if openErr := e.reopen(ctx, topic.Tenant); openErr != nil { diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index f22fb7b0..edb99d97 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -458,6 +458,48 @@ func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) } +// An open that gives up on the ingest stream can leave one behind that +// JetStream goes on to create — in-process, a call fails by timing out — and +// no consumer holds it. A publish goes by the broker's record of the queue, +// not by the stream answering: it opens the queue properly first, consumers +// joined, so its row reaches them rather than a stream nobody reads. +func TestEmbeddedNATS_Publish_OpensAQueueItsOpenGaveUpOn(t *testing.T) { + dir := t.TempDir() + block := filepath.Join(dir, "jetstream", "$G", "streams", ingestStreamName("acme")) + require.NoError(t, os.MkdirAll(filepath.Dir(block), 0o750)) + require.NoError(t, os.WriteFile(block, nil, 0o600)) + e := openEmbedded(t, dir) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer"}) + require.NoError(t, err) + got := make(chan string, 1) + stop, _, err := cons.Consume(func(msg *Message) { + got <- string(msg.Data) + _ = msg.Ack() + }, 10) + require.NoError(t, err) + defer stop() + + require.Error(t, e.SetMaxBytes(ctx, "acme", testBudget), "the ingest stream cannot open") + if err := os.Remove(block); err != nil { + require.ErrorIs(t, err, os.ErrNotExist) + } + // JetStream creates it after all, behind the broker's back. + _, err = e.js.CreateStream(ctx, ingestStreamConfig("acme", testBudget)) + require.NoError(t, err) + + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + select { + case data := <-got: + assert.Equal(t, "x", data) + case <-ctx.Done(): + t.Fatal("the row reached no consumer") + } + assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) +} + // A resize whose dead-letter update fails undoes the ingest one, back to the // cap the ingest stream had. That is not the budget applied in full: a boot // that found the pair split applied none, and a cap of 0 would leave the From 7cdb794d0e946d6f6edc83e1e9142f0191f4f9fe Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 22:02:13 -0400 Subject: [PATCH 009/108] docs(mq): a consumer that cannot join a queue opened at runtime; review fixes --- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 3 ++- docs/src/content/docs/settings-directory.mdx | 2 +- internal/mq/embedded.go | 17 +++++++++-------- 4 files changed, 13 insertions(+), 11 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 9da998d1..c91e7303 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -147,7 +147,7 @@ The **only** package that imports NATS/JetStream — a `depguard` rule in `.gola - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. -- **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. It runs on each tenant's stream at that tenant's cutoff. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. +- **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. - **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 4f7cc1b7..8e57d823 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -33,7 +33,7 @@ WaveHouse does not currently expose a knob to relax this — `SyncAlways` is alw Because the publish blocks on `fsync`, **your typical ingest latency is your storage's typical `fsync` latency, and your worst-case publish is your storage's worst-case `fsync`.** When that tail is healthy (sub-millisecond to single-digit milliseconds) the guarantee is essentially free. When it is not, the same code path that handles every production message stalls: - Publishes block for the duration of the `fsync`, so a multi-second `fsync` tail is a multi-second ingest tail. -- The embedded server's consumer setup and every publish run under the JetStream client's request timeout, and opening or resizing a tenant's queue under a ten-second budget of WaveHouse's own; a slow-enough substrate makes them exceed it. The symptom when a tenant's queue first opens — at the boot or reload that first serves the tenant — is `open dlq stream: ... context deadline exceeded`, or `open ingest stream: ...` (the two share the budget); a boot that finds every queue already at its budget writes nothing, so there the first publish is where it shows. +- The embedded server's consumer setup at boot and every publish run under the JetStream client's request timeout, and opening or resizing a tenant's queue — and joining the consumers to one that opens while the server runs — under ten-second budgets of WaveHouse's own; a slow-enough substrate makes them exceed it. The symptom when a tenant's queue first opens — at the boot or reload that first serves the tenant — is `open dlq stream: ... context deadline exceeded`, or `open ingest stream: ...` (the two share the budget); a boot that finds every queue already at its budget writes nothing, so there the first publish is where it shows. - If the worker cannot drain to ClickHouse faster than producers publish, a tenant's stream fills toward its [`mq.max_bytes_gb`](/settings-directory#message-queue) and the API returns `503` to that tenant ([backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs)). ## Where `SyncAlways` is cheap vs. expensive @@ -94,6 +94,7 @@ A self-contained `wavehouse storage-check` preflight subcommand that bakes this If you see any of these, benchmark the `/nats` volume as above: - `open dlq stream: ... context deadline exceeded`, or `open ingest stream: ...`, when a tenant's queue first opens, at the boot or reload that first serves the tenant. +- `ingest consumer delivery ended; ingestion has stopped` with `join its queue: ... context deadline exceeded`, and the process exiting, when a tenant's queue opens while the server runs and the ingest worker's consumer cannot join it in time; the stream hub's consumer failing the same way logs `a tenant's events do not reach this consumer until the next boot` instead. - Ingest p99 latency in the seconds, or occasional `200`s that take multiple seconds to return. - Intermittent `503 Service Unavailable` from `/v1/ingest` when ClickHouse is healthy (the worker can't drain fast enough because acking is `fsync`-bound). - Flaky CI or load tests that pass on fast storage and fail on a shared/virtualized host. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index d34c0a65..15f4d24e 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -221,7 +221,7 @@ A tenant's dead-letter stream is opened when the tenant is first served (an empt ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. A queue that opens while the server runs but that a consumer cannot join is different: if the ingest worker's cannot, the process exits with the error, and its restart joins the queue at boot; if the stream hub's cannot, that is logged (`a tenant's events do not reach this consumer until the next boot`), and the tenant's streams get no live rows, gap-fill aside, until a restart. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. **Sizing the volume.** Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume — counting every tenant ever served on it, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream (up to a tenth of its last budget). A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index bebcfa63..aa102b85 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -66,10 +66,11 @@ type EmbeddedNATS struct { // consumers are the durable consumers held on every tenant's queue, each // joined to a queue as it opens. consumers []*fanIn - // opened holds the tenants whose queue has both streams and every - // registered consumer joined — what Publish trusts, rather than a stream - // answering: an open that gave up can leave behind a stream JetStream goes - // on to create, which no consumer holds. Written under mu, read without it. + // opened holds the tenants whose queue has both streams, every registered + // consumer joined to it or told it could not be (fanIn.fail) — what + // Publish trusts, rather than a stream answering: an open that gave up can + // leave behind a stream JetStream goes on to create, which no consumer + // holds. Written under mu, read without it. opened sync.Map // tenant.ID → struct{} } @@ -274,10 +275,10 @@ func (e *EmbeddedNATS) ingestTenants() []tenant.ID { } // record brings opened in line with what the broker knows of tenant id's -// queue. Both streams known means every consumer holds the queue too: apply -// joins the consumers to a queue it opens before this records it, and a -// consumer registered later joins every ingest stream there is. Under e.mu -// (or before e is shared). +// queue. Both streams known means every consumer has been joined to the +// queue too, or told it could not be: apply joins the consumers to a queue it +// opens before this records it, and a consumer registered later joins every +// ingest stream there is. Under e.mu (or before e is shared). func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { if q.ingest && q.dlq { e.opened.Store(id, struct{}{}) From e199a035f28dea5b23d2809fdd6f4a1639b4cc07 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 22:31:47 -0400 Subject: [PATCH 010/108] fix(mq): pace publish-side retries of a queue that cannot open; review fixes --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/ingest-pipeline.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/wire.go | 7 ++- internal/mq/embedded.go | 56 ++++++++++++++++-- internal/mq/embedded_test.go | 60 +++++++++++++++++++- 7 files changed, 117 insertions(+), 14 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c24d0df9..27cbd578 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index c91e7303..fb6f03fc 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -148,7 +148,7 @@ The **only** package that imports NATS/JetStream — a `depguard` rule in `.gola - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index 448fa6a0..f314ca49 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -202,7 +202,7 @@ Messages still sitting in `msgChan` or the consumer's prefetch buffer at shutdow ### When the consumer dies -Delivery can end underneath a running worker: the durable consumer is deleted, or the MQ connection closes. The broker client reports that only through an asynchronous error callback and then stops delivering — no message ever arrives to say so, so a loop that only watches `msgChan` would wait forever while the API kept accepting events nothing writes. `mq.Consumer.Consume` therefore returns a `failed` channel next to `stop` (`mq.ErrDeliveryEnded`, wrapping the broker's reason), and `dispatchLoop` selects on it beside `ctx.Done()` and `msgChan`. On a failure it runs the same bottom-up drain as a shutdown — the rows already in hand are flushed and acked, not abandoned — and then reports the error on the worker's own `failed` channel. A consumer that cannot start at all takes the same path. +Delivery can end underneath a running worker: the durable consumer is deleted, the MQ connection closes, or a tenant's queue opened while the server runs cannot be joined. The broker client reports the first two only through an asynchronous error callback and then stops delivering, and `internal/mq` reports the third when it opens the queue — no message ever arrives to say so, so a loop that only watches `msgChan` would wait forever while the API kept accepting events nothing writes. `mq.Consumer.Consume` therefore returns a `failed` channel next to `stop` (`mq.ErrDeliveryEnded`, wrapping the broker's reason), and `dispatchLoop` selects on it beside `ctx.Done()` and `msgChan`. On a failure it runs the same bottom-up drain as a shutdown — the rows already in hand are flushed and acked, not abandoned — and then reports the error on the worker's own `failed` channel. A consumer that cannot start at all takes the same path. The worker does not try to revive the consumer. The app's ingest-worker component returns the error from `app.Run`, which stops every other component and exits non-zero, the same way any failed component does; the supervisor's restart recreates the durable consumer at boot, and everything unacked is redelivered (at-least-once). Passing conditions the client also reports through that callback (a missed heartbeat, a leadership change) are logged at `WARN` and do not end the worker. With the embedded broker (`DontListen`, no external client that could delete a durable) this path is hard to reach; the likeliest way in is a tenant's queue, opened at runtime, that the consumer cannot join. It matters more once a remote broker exists. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 15f4d24e..c1580690 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -221,7 +221,7 @@ A tenant's dead-letter stream is opened when the tenant is first served (an empt ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. A queue that opens while the server runs but that a consumer cannot join is different: if the ingest worker's cannot, the process exits with the error, and its restart joins the queue at boot; if the stream hub's cannot, that is logged (`a tenant's events do not reach this consumer until the next boot`), and the tenant's streams get no live rows, gap-fill aside, until a restart. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each reload trying the queue again, and so does a publish, at most once every five seconds — while every other tenant carries on. A queue that opens while the server runs but that a consumer cannot join is different: if the ingest worker's cannot, the process exits with the error, and its restart joins the queue at boot; if the stream hub's cannot, that is logged (`a tenant's events do not reach this consumer until the next boot`), and the tenant's `GET /v1/stream` connections get no live rows, gap-fill aside, until a restart. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. **Sizing the volume.** Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume — counting every tenant ever served on it, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream (up to a tenth of its last budget). A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. diff --git a/internal/app/wire.go b/internal/app/wire.go index 968d5f34..ac494bea 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -530,9 +530,10 @@ func (a *App) wireDedupe() error { // refuses boot, like every other store, and on a reload logs it, keeping the // previous budget; a nested directory logs it at boot too, so it never costs // the process — the tenant's ingest answers 503 until its queue opens, each -// publish and each reload trying again. The hook is registered before the -// boot apply, as the dedupe one is. The boot apply runs on ctx, New's, so a -// stop signaled during a boot that opens many queues is not held up by them. +// reload trying again, and publishes too at the pace the MQ allows. The hook +// is registered before the boot apply, as the dedupe one is. The boot apply +// runs on ctx, New's, so a stop signaled during a boot that opens many queues +// is not held up by them. func (a *App) wireMQ(ctx context.Context) error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index aa102b85..32d5047d 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -18,6 +18,7 @@ import ( natsserver "github.com/nats-io/nats-server/v2/server" "github.com/nats-io/nats.go" "github.com/nats-io/nats.go/jetstream" + "golang.org/x/sync/singleflight" ) // slogNATSLogger adapts the default slog logger to the natsserver.Logger @@ -72,6 +73,20 @@ type EmbeddedNATS struct { // leave behind a stream JetStream goes on to create, which no consumer // holds. Written under mu, read without it. opened sync.Map // tenant.ID → struct{} + // reopening merges into one attempt the publishes that find the same + // tenant's queue not open, and failedOpen holds, for a tenant whose last + // such attempt failed, its error and until when its publishes take that + // as their answer (openForPublish). + reopening singleflight.Group + failedOpen sync.Map // tenant.ID → openFailure +} + +// openFailure is a publish's failed attempt to open a tenant's queue, and +// until when the tenant's publishes are refused with its error rather than +// trying again. +type openFailure struct { + until time.Time + err error } // tenantQueue is what the broker knows of one tenant's queue. @@ -111,6 +126,9 @@ const ( // resizeTimeouts when it opens a queue: the consumers join on a budget of // their own (apply). rollbackTimeout = 5 * time.Second + // publishRetry is how long a tenant's publishes are refused at once after + // one failed to open its queue (openForPublish). + publishRetry = 5 * time.Second ) // errNoQueue is why a publish or park finds no queue it can open: no budget @@ -282,6 +300,7 @@ func (e *EmbeddedNATS) ingestTenants() []tenant.ID { func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { if q.ingest && q.dlq { e.opened.Store(id, struct{}{}) + e.failedOpen.Delete(id) } else { e.opened.Delete(id) } @@ -492,23 +511,24 @@ func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { // subject). A tenant with no queue has one opened at the budget last asked // for it (see SetMaxBytes) — and so does one whose stream exists but whose // queue the broker has not recorded open, since no consumer may hold that -// stream. A queue that cannot be opened — none asked for yet, or JetStream -// refused it — and a queue at its byte budget (DiscardNew) are reported as -// ErrQueueFull: either way the tenant's queue takes nothing now, and a retry -// is the caller's answer. +// stream (see openForPublish for how often a publish tries). A queue that +// cannot be opened — none asked for yet, or JetStream refused it — and a +// queue at its byte budget (DiscardNew) are reported as ErrQueueFull: either +// way the tenant's queue takes nothing now, and a retry is the caller's +// answer. func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, opts ...PublishOpt) error { subj, err := subject(ingestPrefix, topic) if err != nil { return err } if _, ok := e.opened.Load(topic.Tenant); !ok { - if openErr := e.reopen(ctx, topic.Tenant); openErr != nil { + if openErr := e.openForPublish(ctx, topic.Tenant); openErr != nil { return fmt.Errorf("%w: %w", ErrQueueFull, openErr) } } err = e.publish(ctx, subj, data, opts) if errors.Is(err, jetstream.ErrNoStreamResponse) { - if openErr := e.reopen(ctx, topic.Tenant); openErr != nil { + if openErr := e.openForPublish(ctx, topic.Tenant); openErr != nil { return fmt.Errorf("%w: %w", ErrQueueFull, openErr) } err = e.publish(ctx, subj, data, opts) @@ -521,6 +541,30 @@ func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, op return err } +// openForPublish opens tenant id's queue for a publish that found it not open +// (reopen). The publishes that find it so at the same time share one +// attempt, and after an attempt fails the tenant's publishes get its error at +// once, without taking mu, until publishRetry has passed: under clients +// retrying, a queue that cannot open would otherwise hold mu for attempt +// after attempt, and every other tenant's open, resize and reload waits on +// mu. A reload that applies the tenant's budget retries it regardless +// (SetMaxBytes). +func (e *EmbeddedNATS) openForPublish(ctx context.Context, id tenant.ID) error { + if v, ok := e.failedOpen.Load(id); ok { + if f := v.(openFailure); time.Now().Before(f.until) { + return f.err + } + } + _, err, _ := e.reopening.Do(string(id), func() (any, error) { + err := e.reopen(ctx, id) + if err != nil { + e.failedOpen.Store(id, openFailure{until: time.Now().Add(publishRetry), err: err}) + } + return nil, err + }) + return err +} + // DeadLetter stores msg's data on its topic's dead-letter subject, in its // tenant's queue — the subject it arrived on with the ingest prefix swapped // for the dead-letter one, nothing decoded or re-encoded. The dead-letter diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index edb99d97..b27d1a05 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -424,7 +424,8 @@ func TestEmbeddedNATS_SetMaxBytes_IngestFailureChangesNothing(t *testing.T) { // errors and applies no budget, and a publish is refused as a full queue, // while every other tenant's queue opens after it (which a store limit at // the very top of the int64 range would refuse: see NewEmbedded). Once the -// cause is gone, a publish opens the queue at the budget last asked for it. +// cause is gone, a reload opens the queue at the budget last asked for it, +// however recently a publish tried. func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { dir := t.TempDir() // The dead-letter stream is the first of the pair to open. A failed open @@ -454,10 +455,67 @@ func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { if err := os.Remove(block); err != nil { require.ErrorIs(t, err, os.ErrNotExist) } + require.NoError(t, e.SetMaxBytes(ctx, "acme", testBudget)) require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) } +// After a publish fails to open its tenant's queue, the tenant's publishes +// are refused at once, without waiting on the broker's lock, until +// publishRetry has passed: under clients retrying, one tenant's broken queue +// would otherwise hold the lock that every other tenant's open, resize and +// reload takes. Once the window has passed, a publish tries again. +func TestEmbeddedNATS_Publish_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T) { + dir := t.TempDir() + block := filepath.Join(dir, "jetstream", "$G", "streams", dlqStreamName("acme")) + obstruct := func() { + t.Helper() + require.NoError(t, os.MkdirAll(filepath.Dir(block), 0o750)) + require.NoError(t, os.WriteFile(block, nil, 0o600)) + } + obstruct() + e := openEmbedded(t, dir) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + acme := Topic{Tenant: "acme", Table: "t"} + + require.Error(t, e.SetMaxBytes(ctx, "acme", testBudget)) + obstruct() + require.ErrorIs(t, e.Publish(ctx, acme, []byte("x")), ErrQueueFull, "the publish's own attempt fails") + if err := os.Remove(block); err != nil { + require.ErrorIs(t, err, os.ErrNotExist) + } + + // The queue could open now, but within the window a publish tries + // nothing: it is refused while the lock is held elsewhere. + e.mu.Lock() + var paced error + done := make(chan struct{}) + go func() { + defer close(done) + paced = e.Publish(ctx, acme, []byte("x")) + }() + var returned bool + select { + case <-done: + returned = true + case <-time.After(2 * time.Second): + } + e.mu.Unlock() + <-done + require.True(t, returned, "a paced publish waited on the broker's lock") + require.ErrorIs(t, paced, ErrQueueFull) + assert.Zero(t, e.MaxBytes("acme")) + + v, ok := e.failedOpen.Load(tenant.ID("acme")) + require.True(t, ok) + failed := v.(openFailure) + failed.until = time.Now() + e.failedOpen.Store(tenant.ID("acme"), failed) + require.NoError(t, e.Publish(ctx, acme, []byte("x")), "once the window has passed") + assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) +} + // An open that gives up on the ingest stream can leave one behind that // JetStream goes on to create — in-process, a call fails by timing out — and // no consumer holds it. A publish goes by the broker's record of the queue, From 55137d8ec23ff7278323f119a40ef9e9ec17f9e0 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 22:55:33 -0400 Subject: [PATCH 011/108] test(mq): one tenant's failed purge stops no other; review fixes --- docs/src/content/docs/sdk/streaming.md | 2 +- internal/mq/embedded_test.go | 110 +++++++++++++++++++++++++ 2 files changed, 111 insertions(+), 1 deletion(-) diff --git a/docs/src/content/docs/sdk/streaming.md b/docs/src/content/docs/sdk/streaming.md index 94a7e334..27ee3a95 100644 --- a/docs/src/content/docs/sdk/streaming.md +++ b/docs/src/content/docs/sdk/streaming.md @@ -116,7 +116,7 @@ A dropped stream reconnects on a jittered exponential backoff, capped at 30s, an :::caution[Resumption is at-least-once, and time-bounded] Delivery across a reconnect is **at-least-once**. The `Last-Event-ID` the client sends is the last event's `received_timestamp`, and the server replays from that instant *inclusively* — so the last event you already saw, and anything sharing its timestamp, arrives again. The SDK does not deduplicate live frames — `liveQuery()` makes one pass at the backfill seam, and only under an ascending order ([#449](https://github.com/Wave-RF/WaveHouse/issues/449)) — so key on `timestamp` plus your own row identity if duplicates matter. -Replay is also bounded by the server's [`stream.gap_window_minutes`](/settings-directory#streaming) — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. The same silence applies across a server upgrade to this release: the server deletes the previous release's queue at boot, so a replay spanning the upgrade omits the events published before it, without an error — backfill over REST if you need them. +Replay is also bounded by the [`stream.gap_window_minutes`](/settings-directory#streaming) of the tenant you stream from — 15 minutes by default. A drop longer than that resumes with a hole and no signal, because the purged messages are simply gone. The same silence applies to a replay spanning the upgrade across the v2 ingest envelope, whose boot deletes the earlier build's queue — see [Upgrading across the v2 ingest envelope](/deployment#upgrading-across-the-v2-ingest-envelope); backfill over REST if you need those events. **A column-set change across a gap-fill is a known limitation.** If the table's columns change while you are connected *and* your client replays across that change, live rows arriving after the replay may not be preceded by a fresh `event: schema` frame until the columns next change or you reconnect. The SDK drops a row whose **length** disagrees with the list it was last told, rather than zipping it under the wrong names — so an added or removed column costs you rows, not wrong ones. A **same-length** change is the residual case the arity check cannot see: a `RENAME COLUMN`, or a drop paired with an add, zips values under the wrong names until the next announcement. Reconnecting resynchronizes either way. Full schema-change handling is deferred to the schema-versioning work ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)). ::: diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index b27d1a05..85fd5719 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -1000,6 +1000,56 @@ func TestEmbeddedNATS_PurgeAcked_EachTenantAtItsOwnCutoff(t *testing.T) { assert.Zero(t, msgs("initech"), "a tenant the cutoffs do not name keeps nothing it has acknowledged") } +// One tenant's purge failing stops no other tenant's: the errors say which +// failed, and the sweep goes on to the next tenant at its own cutoff — here +// after one whose durable is gone and one whose stream is. A sweep whose +// context has already ended touches no tenant. +func TestEmbeddedNATS_PurgeAcked_OneTenantsFailureStopsNoOther(t *testing.T) { + ids := []tenant.ID{"acme", "globex", "initech", "umbrella"} + e := newTestEmbedded(t, ids...) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + for _, id := range ids { + for i := range 2 { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "p"}, []byte{byte(i)})) + } + } + ackAll(t, e, "buffer", 8) + for _, id := range ids { + s, err := e.stream(ctx, ingestStreamName(id)) + require.NoError(t, err) + require.Eventually(t, func() bool { + floor, err := s.consumerAckFloor(ctx, "buffer") + return err == nil && floor == 2 + }, 5*time.Second, 20*time.Millisecond, id) + } + msgs := func(id tenant.ID) uint64 { + s, err := e.stream(ctx, ingestStreamName(id)) + require.NoError(t, err) + st, err := s.state(ctx, "") + require.NoError(t, err) + return st.Msgs + } + + ended, end := context.WithCancel(ctx) + end() + purged, err := e.PurgeAcked(ended, "buffer", nil) + require.ErrorIs(t, err, context.Canceled) + assert.False(t, purged) + assert.Equal(t, uint64(2), msgs("umbrella"), "a sweep whose context has ended touches nothing") + + require.NoError(t, e.js.DeleteConsumer(ctx, ingestStreamName("acme"), "buffer")) + require.NoError(t, e.js.DeleteStream(ctx, ingestStreamName("globex"))) + purged, err = e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{"initech": time.Now().Add(-time.Hour)}) + require.ErrorIs(t, err, ErrConsumerNotFound, "acme's durable is gone") + require.ErrorContains(t, err, "tenant globex: get stream") + assert.True(t, purged, "the tenants after them are purged all the same") + assert.Equal(t, uint64(2), msgs("acme")) + assert.Equal(t, uint64(2), msgs("initech"), "kept for its own window") + assert.Zero(t, msgs("umbrella")) +} + // The isolation per-tenant queues buy: a tenant at MaxAckPending, or one // whose handler is stuck, holds back its own delivery and no other tenant's — // each tenant's messages arrive on a delivery of their own, in order. @@ -1328,3 +1378,63 @@ func TestNewEmbedded_TakesStockOfTheQueuesOnDisk(t *testing.T) { require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) assert.Equal(t, int64(8<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) } + +// A durable found on disk is kept as it stands when it holds the settings +// asked for — a boot over many queues writes nothing it need not — and is +// updated in place when they differ; either way delivery resumes past what it +// acknowledged before the restart. +func TestEmbeddedNATS_ADurableOnDiskIsReusedAcrossARestart(t *testing.T) { + for _, tt := range []struct { + name string + maxAckPending int + }{ + {"same settings", 10}, + {"other settings", 20}, + } { + t.Run(tt.name, func(t *testing.T) { + dir := t.TempDir() + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + topic := Topic{Tenant: "acme", Table: "t"} + + first, err := NewEmbedded(dir) + require.NoError(t, err) + require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) + require.NoError(t, first.Publish(ctx, topic, []byte{0})) + cons, err := first.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 10}) + require.NoError(t, err) + acked := make(chan error, 2) + stop, _, err := cons.Consume(func(msg *Message) { acked <- msg.DoubleAck(ctx) }, 1) + require.NoError(t, err) + select { + case err := <-acked: + require.NoError(t, err) + case <-ctx.Done(): + t.Fatal("the first row was not delivered") + } + stop() + require.NoError(t, first.Close()) + + e := openEmbedded(t, dir) + require.NoError(t, e.Publish(ctx, topic, []byte{1})) + cons, err = e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: tt.maxAckPending}) + require.NoError(t, err) + got := make(chan byte, 2) + stop, _, err = cons.Consume(func(msg *Message) { + _ = msg.Ack() + got <- msg.Data[0] + }, 4) + require.NoError(t, err) + t.Cleanup(stop) + select { + case b := <-got: + assert.Equal(t, byte(1), b, "delivery resumes past what was acknowledged before the restart") + case <-ctx.Done(): + t.Fatal("the row published after the restart was not delivered") + } + c, err := e.js.Consumer(ctx, ingestStreamName("acme"), "buffer") + require.NoError(t, err) + assert.Equal(t, tt.maxAckPending, c.CachedInfo().Config.MaxAckPending) + }) + } +} From 060ca9ea040984864da01ec553a1433109208523 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 23:17:19 -0400 Subject: [PATCH 012/108] docs(mq): size a dead-letter stream the shrink guard kept; review fixes --- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index df65cfe4..6fc71291 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -745,7 +745,7 @@ Triggers an immediate re-discovery of the `?tenant=`'s ClickHouse table schemas #### `GET /v1/ops/dlq/stats` — DLQ Statistics -Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, since its queue is kept (nothing deletes it). The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream is opened when the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. +Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The tenant is looked up in the message queue, not the settings, so a tenant whose folder was rejected or removed is read like one being served, since its queue is kept (nothing deletes it). The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream is opened when the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. **Error responses:** diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c1580690..93bfdd6a 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -223,7 +223,7 @@ A tenant's dead-letter stream is opened when the tenant is first served (an empt - `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the Active Sweeper purges it back under the limit — what it purges is what is both written to ClickHouse and past the tenant's `stream.gap_window_minutes`, and nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each reload trying the queue again, and so does a publish, at most once every five seconds — while every other tenant carries on. A queue that opens while the server runs but that a consumer cannot join is different: if the ingest worker's cannot, the process exits with the error, and its restart joins the queue at boot; if the stream hub's cannot, that is logged (`a tenant's events do not reach this consumer until the next boot`), and the tenant's `GET /v1/stream` connections get no live rows, gap-fill aside, until a restart. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. -**Sizing the volume.** Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep the tenants' budgets, plus a tenth of each for their dead-letter streams, within the free space of the `/nats` volume — counting every tenant ever served on it, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream (up to a tenth of its last budget). A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. +**Sizing the volume.** Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to ([#138](https://github.com/Wave-RF/WaveHouse/issues/138)) — so keep within the free space of the `/nats` volume every tenant's budget plus its dead-letter stream's cap: a tenth of the budget, or what the stream held when a smaller budget arrived, if that is more. Count every tenant ever served on the volume, not only those served now: a rejected or removed tenant's queue is kept and nothing deletes it, so what it holds goes on holding disk — a rejected tenant's replay history, and the rows parked on either one's dead-letter stream, up to that cap. A disk that fills before a budget does fails every tenant's writes, not one: the failed write is logged, the publish goes unanswered until it times out, and ingest answers `500` (`publish failed`) for every tenant on that volume, not the `503` with `Retry-After` of a full budget. It does not clear on its own: a stream that failed a write refuses every later one until WaveHouse restarts, so free the space and then restart. ## Streaming From 3022b926392f9cdfa16931ada447607b94fe09b7 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:27:32 -0400 Subject: [PATCH 013/108] feat(config): choose each layer's implementation at boot mq.backend, cache.backend, dedupe.backend and coord.backend select each layer's implementation; only today's in-process one exists per layer and it is the default. Validate refuses an unknown value, internal/app picks the implementation in one switch per layer, data_dir is probed only when a selected backend keeps state there, and boot logs Config.Warnings. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 1 + cmd/wavehouse/main.go | 19 ++- config.yaml | 10 ++ docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 31 +++- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app.go | 4 + internal/app/app_test.go | 27 +++- internal/app/wire.go | 63 ++++++-- internal/config/backends.go | 143 ++++++++++++++++ internal/config/backends_test.go | 161 +++++++++++++++++++ internal/config/config.go | 12 +- internal/config/config_test.go | 10 +- tests/integration/setup_test.go | 5 +- tests/integration/tenants_test.go | 5 +- 15 files changed, 453 insertions(+), 42 deletions(-) create mode 100644 internal/config/backends.go create mode 100644 internal/config/backends_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c715..9f2e821b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/cmd/wavehouse/main.go b/cmd/wavehouse/main.go index 042257e2..3a1304f0 100644 --- a/cmd/wavehouse/main.go +++ b/cmd/wavehouse/main.go @@ -171,14 +171,17 @@ func run(ctx context.Context) int { return 1 } - // data_dir must be writable before anything dials out, so the refusal - // (and, for the typical cause — a bind mount owned by root rather than - // UID 65532 — the remediation) lands at the top of the log rather than - // after ClickHouse discovery. NATS and Pebble still fail loud on their - // own if the directory changes underneath us. - if err := config.CheckDataDir(cfg.DataDir); err != nil { - logger.Error("check data_dir", "error", err) - return 1 + // data_dir, when a selected backend keeps state there, must be writable + // before anything dials out, so the refusal (and, for the typical cause — + // a bind mount owned by root rather than UID 65532 — the remediation) + // lands at the top of the log rather than after ClickHouse discovery. + // NATS and Pebble still fail loud on their own if the directory changes + // underneath us. + if cfg.NeedsDataDir() { + if err := config.CheckDataDir(cfg.DataDir); err != nil { + logger.Error("check data_dir", "error", err) + return 1 + } } a, err := app.New(ctx, app.Options{ diff --git a/config.yaml b/config.yaml index 53a43502..65c384ea 100644 --- a/config.yaml +++ b/config.yaml @@ -43,9 +43,19 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none +# Each layer's implementation, chosen at boot. Only the in-process backend +# exists for each today, and it is the default. +mq: + backend: embedded # NATS JetStream under /nats +dedupe: + backend: pebble # Pebble under /pebble +coord: + backend: local + # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. cache: + backend: local l1_max_cost: 67108864 # Auth has no on/off switch — the JWT middleware always runs. A request with no diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54..938751ff 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8a709426..8d311adc 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -17,7 +17,7 @@ WaveHouse is configured via a YAML file with environment variable overrides. All 2. Environment variables override any values from the YAML file. 3. If no config file exists, all values are read from environment variables. Every key has a default except `settings.dir` (`WH_SETTINGS_DIR`), which must be set either way. 4. Both sources are **strict**. A YAML key this page doesn't list — a typo, or a tunable that has moved to the settings directory (`dlq.enabled`, `clickhouse.addr`, `stream.*`, a leftover `policy:` or `pipes:` block, …) — refuses to boot and names every offending key, so nothing is read, ignored, and believed. A `WH_*` environment variable that binds to no key on this page (`WH_DEDUPE_ENABLED`, `WH_CH_ADDR`, a misspelling) refuses to boot the same way. Two variables have no YAML key and are exempt because they are not config keys at all but process-level settings `main` reads directly: `WH_CONFIG` (below), which locates the file, and `WH_LOG_LEVEL`. Only the `WH_` prefix is checked, since the environment always carries names that aren't WaveHouse's. One outside source does share the prefix. Kubernetes injects `{SERVICE}_SERVICE_HOST`, `{SERVICE}_PORT`, and similar link variables into every pod in a Service's own namespace, for each Service with a cluster IP that existed before the pod started (a headless Service injects nothing, and a Service in another namespace is harmless). The name is uppercased with `-` mapped to `_`, so a Service named `wh` produces `WH_SERVICE_HOST` and `WH_PORT`, one named `wh-foo` produces `WH_FOO_SERVICE_HOST` and `WH_FOO_PORT`, and either way the pod refuses to boot on its next restart. Set `enableServiceLinks: false` on the pod spec, or name the Service something else. The error says so. -5. Before anything dials out, `data_dir` is probed, and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. +5. Before anything dials out, `data_dir` is probed — when a selected [backend](#backends) keeps state there, as the in-process `mq` and `dedupe` backends do — and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. Boot is the validator for this half of configuration: there is no dry run, and a refused boot with the offending key, variable, or path named in the error is the loud signal. The hot-reloadable half has a dry run — `wavehouse validate` — because it is edited under a running server; boot config only ever takes effect through a restart, so the restart is where it is checked. @@ -37,6 +37,19 @@ This page is boot config only — what the platform operator owns (wiring, lifec | --- | --- | ------- | ----------- | | `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble, holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | +### Backends + +Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | +| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | +| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | + +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. + ### Server | YAML Key | Env Var | Default | Description | @@ -97,6 +110,8 @@ WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so ### Message Queue (NATS) +This section describes the `embedded` [backend](#backends), the only one today. + Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. **Durability.** The embedded server runs with JetStream `SyncAlways`, so every event is `fsync`'d to disk before `POST /v1/ingest` returns `200`. This makes your storage's `fsync` latency your ingest latency floor — see [Durability & Storage](/durability) to check whether your substrate can sustain it. There is no knob to relax this today ([#139](https://github.com/Wave-RF/WaveHouse/issues/139) tracks a configurable group-commit interval). @@ -191,9 +206,19 @@ clickhouse: # headers and pool sizes are settings (config.json) max_total_conns: 0 # ceiling on open native connections; 0 = none +mq: + backend: embedded # in-process NATS JetStream under /nats + cache: + backend: local l1_max_cost: 67108864 +dedupe: + backend: pebble # in-process Pebble under /pebble + +coord: + backend: local + auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) operator_key: "" # non-JWT full-access operator credential (Authorization: Operator , or X-Operator-Key); empty disables @@ -239,7 +264,11 @@ WH_SERVER_SHUTDOWN_TIMEOUT=10 WH_CH_PASSWORD= WH_CH_MAX_TOTAL_CONNS=0 +WH_MQ_BACKEND=embedded +WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 +WH_DEDUPE_BACKEND=pebble +WH_COORD_BACKEND=local WH_AUTH_JWT_SECRET=change-me-in-production WH_AUTH_OPERATOR_KEY= diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e7a19b..56135462 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -179,7 +179,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication diff --git a/internal/app/app.go b/internal/app/app.go index dc15bf5d..726dad19 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -170,6 +170,10 @@ func New(ctx context.Context, opts Options) (app *App, err error) { return nil, err } a.wireObservability(ctx) + // After observability, so an OTLP log pipeline carries them too. + for _, w := range a.cfg.Warnings() { + slog.Warn(w) + } if err := a.wireClickHouse(); err != nil { return nil, err } diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d9..10658ddc 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -95,7 +95,10 @@ func testConfig(t *testing.T, settingsDir string) *config.Config { return &config.Config{ DataDir: t.TempDir(), Server: config.Server{Port: closedPort(t), ShutdownTimeout: 2}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Auth: config.Auth{JWTSecret: "unit-test-secret"}, Settings: config.Settings{Dir: settingsDir}, } @@ -535,6 +538,28 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { assert.False(t, restored.Open(), "Close releases every open store") } +// Validate refuses a backend no layer has a case for, so the switch's default +// is reached only by a Config built by hand; it must refuse boot, not wire +// nothing. +func TestNew_RefusesALayerWithoutABackend(t *testing.T) { + for _, tc := range []struct { + key string + unset func(*config.Config) + }{ + {"dedupe.backend", func(c *config.Config) { c.Dedupe.Backend = "" }}, + {"mq.backend", func(c *config.Config) { c.MQ.Backend = "" }}, + {"cache.backend", func(c *config.Config) { c.Cache.Backend = "" }}, + } { + t.Run(tc.key, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, nil)) + tc.unset(cfg) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, tc.key+` "" has no wiring`) + }) + } +} + // A Pebble instance that cannot open follows the registry's own rule for the // shape: a flat directory refuses boot, like every other store, and a nested // one fails closed for every tenant with dedupe on, since they share the diff --git a/internal/app/wire.go b/internal/app/wire.go index 60cbdcd8..d5372025 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -471,7 +471,18 @@ func (a *App) wireDiscovery(ctx context.Context) { a.add(component{name: "schema discovery", close: d.close}) } -// wireDedupe builds the dedupe stores: one per tenant (#583 story 7), each +// wireDedupe builds the dedupe stores — the one place the implementation is +// chosen. +func (a *App) wireDedupe() error { + switch b := a.cfg.Dedupe.Backend; b { + case config.DedupePebble: + return a.wirePebbleDedupe() + default: + return unreachableBackend("dedupe.backend", b) + } +} + +// wirePebbleDedupe builds the dedupe stores: one per tenant (#583 story 7), each // following its own tenant's hot-reloadable dedupe.enabled, over the // embedded Pebble implementation, which is handed data_dir and decides the // rest: every tenant's seen ids in one instance there, open while any @@ -488,7 +499,7 @@ func (a *App) wireDiscovery(ctx context.Context) { // than silently publishing un-deduped, since the files asked for dedupe; // nested fails closed the same way at boot too, for every tenant with // dedupe on, the next reload retrying, so it never costs the process. -func (a *App) wireDedupe() error { +func (a *App) wirePebbleDedupe() error { nested := a.tenants.Nested() embedded := dedupe.NewEmbedded(a.cfg.DataDir) stores := dedupe.NewStores(embedded.Tenant) @@ -530,10 +541,20 @@ func (a *App) wireDedupe() error { return nil } -// wireMQ starts the MQ — the embedded NATS under data_dir/nats, the one -// place the implementation is chosen; everything after it sees mq.Broker — -// and hands it each served tenant's mq.max_bytes_gb, which opens that -// tenant's queue the first time. The budget is hot-reloadable: after every +// wireMQ starts the MQ — the one place the implementation is chosen; +// everything after it sees mq.Broker. +func (a *App) wireMQ() error { + switch b := a.cfg.MQ.Backend; b { + case config.MQEmbedded: + return a.wireEmbeddedMQ() + default: + return unreachableBackend("mq.backend", b) + } +} + +// wireEmbeddedMQ starts the embedded NATS under data_dir/nats and hands it +// each served tenant's mq.max_bytes_gb, which opens that tenant's queue the +// first time. The budget is hot-reloadable: after every // reload the registry applies, each served tenant's is handed over again, // and the MQ owns how it is split across the tenant's queues and keeps them // consistent (see mq.Broker.SetMaxBytes). A tenant no longer served keeps @@ -543,7 +564,7 @@ func (a *App) wireDedupe() error { // previous budget; a nested directory logs it at boot too, so it never costs // the process — the tenant's ingest answers 503 until a reload opens its // queue. The hook is registered before the boot apply, as the dedupe one is. -func (a *App) wireMQ() error { +func (a *App) wireEmbeddedMQ() error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker @@ -591,16 +612,28 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// wireCache opens the query-result cache — the one place the implementation +// is chosen. func (a *App) wireCache() error { - l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) - if err != nil { - return fmt.Errorf("cache init: %w", err) + switch b := a.cfg.Cache.Backend; b { + case config.CacheLocal: + l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + a.cache = l1 + a.add(component{name: "cache", close: withoutContext(l1.Close)}) + return nil + default: + return unreachableBackend("cache.backend", b) } - // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc - a.cache = l1 - a.add(component{name: "cache", close: withoutContext(l1.Close)}) - return nil +} + +// unreachableBackend is each layer switch's default case. config.Validate +// refuses a backend with no case, so reaching it means a Config built by hand +// without one (the zero value is not the default), or a case missing here. +func unreachableBackend[T ~string](key string, got T) error { + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go new file mode 100644 index 00000000..8328cea6 --- /dev/null +++ b/internal/config/backends.go @@ -0,0 +1,143 @@ +package config + +import ( + "fmt" + "slices" + "strings" +) + +// Each layer's implementation is chosen here, once, at boot: `.backend` +// names it, and the default is today's in-process one. Settings for one +// backend go in `.`, a sub-block read only when that backend +// is selected. Adding a backend is its constant in the layer's list, a case +// in the layer's validate for its sub-block, and a case in the layer's +// wire function in internal/app — nothing else in Validate changes. + +// MQBackend names the message queue implementation. +type MQBackend string + +// MQEmbedded is the NATS JetStream server inside this process, under +// /nats. +const MQEmbedded MQBackend = "embedded" + +var mqBackends = []MQBackend{MQEmbedded} + +// MQ selects the message queue. The per-tenant byte budget, mq.max_bytes_gb, +// is a settings-directory key, not this block's. +type MQ struct { + Backend MQBackend `yaml:"backend" env:"WH_MQ_BACKEND" env-default:"embedded"` +} + +func (m MQ) validate() error { + return checkBackend("mq.backend", "WH_MQ_BACKEND", m.Backend, mqBackends) +} + +// CacheBackend names the query-result cache implementation. +type CacheBackend string + +// CacheLocal is the in-process Ristretto cache, sized by cache.l1_max_cost. +const CacheLocal CacheBackend = "local" + +var cacheBackends = []CacheBackend{CacheLocal} + +// Cache selects and sizes the query-result cache. The time-range bucket +// structured queries normalize to is a settings-directory key +// (query.timestamp_bucket_seconds) — query shaping, not process memory. +type Cache struct { + Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` + L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` +} + +func (c Cache) validate() error { + return checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends) +} + +// DedupeBackend names where ingest dedupe keeps the ids it has seen. +type DedupeBackend string + +// DedupePebble is the Pebble instance inside this process, under +// /pebble, opened while any tenant has dedupe on. +const DedupePebble DedupeBackend = "pebble" + +var dedupeBackends = []DedupeBackend{DedupePebble} + +// Dedupe selects the dedupe store. Whether a tenant dedupes, and on which +// field, are settings-directory keys, not this block's. +type Dedupe struct { + Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND" env-default:"pebble"` +} + +func (d Dedupe) validate() error { + return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) +} + +// CoordBackend names the lease implementation singleton work (the sweeper) +// is elected through. +type CoordBackend string + +// CoordLocal holds leases in this process: correct while no other process +// shares its queue. +const CoordLocal CoordBackend = "local" + +var coordBackends = []CoordBackend{CoordLocal} + +// Coord selects the coordination layer. +type Coord struct { + Backend CoordBackend `yaml:"backend" env:"WH_COORD_BACKEND" env-default:"local"` +} + +func (c Coord) validate() error { + return checkBackend("coord.backend", "WH_COORD_BACKEND", c.Backend, coordBackends) +} + +// checkBackend refuses a backend this build has no implementation for, +// listing the ones it has. env repeats the struct tag's literal: a tag can't +// reference a constant. +func checkBackend[T ~string](key, env string, got T, valid []T) error { + if slices.Contains(valid, got) { + return nil + } + names := make([]string, len(valid)) + for i, v := range valid { + names[i] = string(v) + } + return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) +} + +// validateBackends checks every layer's backend and its sub-block. +func (c *Config) validateBackends() error { + for _, check := range []func() error{c.MQ.validate, c.Cache.validate, c.Dedupe.validate, c.Coord.validate} { + if err := check(); err != nil { + return err + } + } + return nil +} + +// Distributed reports whether the message queue is shared with other +// processes. The embedded one listens on no port, so while it is selected +// every process is an island: nothing else can reach its queue. +func (c *Config) Distributed() bool { return c.MQ.Backend != MQEmbedded } + +// NeedsDataDir reports whether a selected backend keeps state under data_dir, +// and so whether boot must probe it (CheckDataDir). +func (c *Config) NeedsDataDir() bool { + return c.MQ.Backend == MQEmbedded || c.Dedupe.Backend == DedupePebble +} + +// Warnings returns what a valid configuration is still likely to get wrong, +// one line each, for boot to log at WARN. They are not errors because each is +// correct for a single replica, and one process cannot count its replicas. +func (c *Config) Warnings() []string { + if !c.Distributed() { + return nil + } + var out []string + if c.Cache.Backend == CacheLocal { + out = append(out, "cache.backend=local with a shared mq.backend is correct for one replica only: an event ingested on another replica never invalidates this one's cache, so its reads stay stale until the cached entry expires") + } + if c.Dedupe.Backend == DedupePebble { + out = append(out, "dedupe.backend=pebble with a shared mq.backend dedupes per replica only: an id seen by another replica is not seen by this one") + } + return out +} diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go new file mode 100644 index 00000000..0e70dc7a --- /dev/null +++ b/internal/config/backends_test.go @@ -0,0 +1,161 @@ +package config + +import ( + "os" + "path/filepath" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// withDefaultBackends sets what Load's env-defaults would: a literal Config +// names no backend, and Validate refuses that. +func withDefaultBackends(c Config) *Config { + c.MQ.Backend, c.Cache.Backend = MQEmbedded, CacheLocal + c.Dedupe.Backend, c.Coord.Backend = DedupePebble, CoordLocal + return &c +} + +func defaultBackends() Config { + return *withDefaultBackends(Config{Server: Server{Port: 8080}, Settings: Settings{Dir: "./settings"}}) +} + +func TestLoad_BackendDefaults(t *testing.T) { + t.Parallel() + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) + assert.False(t, cfg.Distributed()) + assert.True(t, cfg.NeedsDataDir()) + assert.Empty(t, cfg.Warnings()) +} + +func TestLoad_BackendsFromEnv(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "embedded") + t.Setenv("WH_CACHE_BACKEND", "local") + t.Setenv("WH_DEDUPE_BACKEND", "pebble") + t.Setenv("WH_COORD_BACKEND", "local") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +func TestLoad_BackendFromEnvRefusesAnUnknownValue(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "nats") + _, err := Load("nonexistent.yaml") + require.Error(t, err) + assert.Contains(t, err.Error(), `mq.backend (WH_MQ_BACKEND) "nats" is not a backend this build has; valid: embedded`) +} + +func TestLoad_BackendsFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded +cache: + backend: local + l1_max_cost: 1024 +dedupe: + backend: pebble +coord: + backend: local +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, int64(1024), cfg.Cache.L1MaxCost) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +// A sub-block written before its backend exists, and a settings-directory +// key under a block both files share, are unknown keys — not read and ignored. +func TestLoad_BackendBlocksRefuseUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded + max_bytes_gb: 5 + nats: + urls: nats://localhost:4222 +dedupe: + enabled: true +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "dedupe.enabled, mq.max_bytes_gb, mq.nats") + assert.Contains(t, err.Error(), EnvSettingsDir) +} + +func TestUnboundEnv_KnowsTheBackendVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_MQ_BACKEND=embedded", "WH_CACHE_BACKEND=local", + "WH_DEDUPE_BACKEND=pebble", "WH_COORD_BACKEND=local", + })) +} + +func TestValidate_UnknownBackend(t *testing.T) { + t.Parallel() + cases := []struct { + name string + set func(*Config) + want string + }{ + {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, + {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, + {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, + {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, + // The zero value, which a Config built without Load carries. + {"empty", func(c *Config) { c.MQ.Backend = "" }, `mq.backend (WH_MQ_BACKEND) "" is not a backend`}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + require.NoError(t, cfg.Validate()) + tc.set(&cfg) + err := cfg.Validate() + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +// Every warning keys on a shared queue, which no backend offers yet, so the +// value is set directly: Warnings reads the choice, it doesn't validate it. +func TestWarnings_SharedQueue(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.Empty(t, cfg.Warnings()) + + cfg.MQ.Backend = "shared" + require.True(t, cfg.Distributed()) + got := cfg.Warnings() + require.Len(t, got, 2) + assert.Contains(t, got[0], "cache.backend=local") + assert.Contains(t, got[1], "dedupe.backend=pebble") + + cfg.Cache.Backend, cfg.Dedupe.Backend = "shared", "shared" + assert.Empty(t, cfg.Warnings()) +} + +func TestNeedsDataDir(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.True(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = "shared" + assert.True(t, cfg.NeedsDataDir(), "pebble dedupe still keeps state under data_dir") + cfg.Dedupe.Backend = "shared" + assert.False(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = MQEmbedded + assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") +} diff --git a/internal/config/config.go b/internal/config/config.go index 68b0314b..cc756696 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -19,7 +19,10 @@ type Config struct { DataDir string `yaml:"data_dir" env:"WH_DATA_DIR" env-default:"./data"` Server Server `yaml:"server"` ClickHouse ClickHouse `yaml:"clickhouse"` + MQ MQ `yaml:"mq"` Cache Cache `yaml:"cache"` + Dedupe Dedupe `yaml:"dedupe"` + Coord Coord `yaml:"coord"` Auth Auth `yaml:"auth"` OTel OTel `yaml:"otel"` Prometheus Prometheus `yaml:"prometheus"` @@ -132,13 +135,6 @@ type ClickHouse struct { MaxTotalConns int `yaml:"max_total_conns" env:"WH_CH_MAX_TOTAL_CONNS" env-default:"0"` } -// Cache sizes the in-process L1 cache. The time-range bucket structured -// queries normalize to is a settings-directory key -// (query.timestamp_bucket_seconds) — query shaping, not process memory. -type Cache struct { - L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` -} - // Auth holds the authentication secrets. The verifier wiring — `jwks_url`, // `role_claim` — is the settings directory's `auth` block (hot-reloadable: // a change rebuilds the verifier). There is no on/off switch: the middleware always runs. A request @@ -224,7 +220,7 @@ func (c *Config) Validate() error { } } - return nil + return c.validateBackends() } // Load reads config from a YAML file (if it exists) with env var overrides. diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 822d639d..ee8b07cc 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -203,7 +203,7 @@ func TestValidate_SampleRatesIgnoredWhenObservabilityDisabled(t *testing.T) { Logs: OTelLogs{SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { @@ -219,7 +219,7 @@ func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { Logs: OTelLogs{Enabled: false, SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestLoad_Defaults_PrometheusDisabled(t *testing.T) { @@ -336,7 +336,7 @@ func TestValidate_PrometheusV1PathAllowedOnSidecarPort(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/v1/metrics", Port: 9091}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { @@ -348,7 +348,7 @@ func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/metrics", Port: 0}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { @@ -365,7 +365,7 @@ func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { Port: 8080, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } // TestEnvSettingsDir_MatchesStructTag pins the exported constant to the diff --git a/tests/integration/setup_test.go b/tests/integration/setup_test.go index a01064a2..ade560f7 100644 --- a/tests/integration/setup_test.go +++ b/tests/integration/setup_test.go @@ -161,7 +161,10 @@ func setup() (int, func()) { DataDir: dataDir, Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, - Cache: config.Cache{L1MaxCost: 1 << 30}, // 1 GB + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 30}, // 1 GB + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: settingsDir}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) diff --git a/tests/integration/tenants_test.go b/tests/integration/tenants_test.go index 40c5a700..ca00f42f 100644 --- a/tests/integration/tenants_test.go +++ b/tests/integration/tenants_test.go @@ -62,7 +62,10 @@ func TestNestedDirectory_PerTenantPoolsAndDiscovery(t *testing.T) { Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, Auth: config.Auth{OperatorKey: operatorKey}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: root}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) From fde17ba4578a2a03521e085a77fd2d3326a75c7d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:30:43 -0400 Subject: [PATCH 014/108] docs(config): say coord.backend is reserved; sync the boot-config lists Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 7 ++++--- docs/src/content/docs/configuration.mdx | 4 ++-- internal/app/wire.go | 2 +- internal/config/backends.go | 8 ++++---- internal/settings/settings.go | 8 +++++--- 7 files changed, 18 insertions(+), 15 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9f2e821b..6564775a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/config.yaml b/config.yaml index 65c384ea..5519b78c 100644 --- a/config.yaml +++ b/config.yaml @@ -50,7 +50,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 938751ff..42591807 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -89,7 +89,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring -- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. +- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. - **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -115,8 +115,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `config/` — Configuration -- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). -- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load`, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). +- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8d311adc..ee264bf4 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -46,7 +46,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | | `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | -| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. @@ -217,7 +217,7 @@ dedupe: backend: pebble # in-process Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) diff --git a/internal/app/wire.go b/internal/app/wire.go index d5372025..3cfcad4b 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -633,7 +633,7 @@ func (a *App) wireCache() error { // refuses a backend with no case, so reaching it means a Config built by hand // without one (the zero value is not the default), or a case missing here. func unreachableBackend[T ~string](key string, got T) error { - return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name the backend of every layer it wires", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go index 8328cea6..f2ab9330 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -71,12 +71,12 @@ func (d Dedupe) validate() error { return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) } -// CoordBackend names the lease implementation singleton work (the sweeper) -// is elected through. +// CoordBackend names where leases for singleton work (the sweeper) are held. +// Nothing reads it yet: the lease layer (#613) wires it. type CoordBackend string -// CoordLocal holds leases in this process: correct while no other process -// shares its queue. +// CoordLocal holds leases in this process, which is enough while no other +// process shares its queue. const CoordLocal CoordBackend = "local" var coordBackends = []CoordBackend{CoordLocal} diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 55ec089d..8db981b3 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -59,9 +59,11 @@ type PipesFile struct { // TenantConfig is the shape of config.json: the behavioral tunables that // migrate out of boot config. Boot config (config.yaml/env) keeps only what -// cannot change under a running process — resource sizing (`data_dir`, -// `cache.l1_max_cost`), listeners, the observability -// exporters — and the secrets (`clickhouse.password`, `auth.jwt_secret`, +// cannot change under a running process — the implementation each layer +// runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, +// `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, +// `clickhouse.max_total_conns`), listeners, the observability exporters — +// and the secrets (`clickhouse.password`, `auth.jwt_secret`, // `auth.operator_key`), which never belong in a tracked JSON file. Every // block and every top-level key inside it is REQUIRED: the binary carries no // compiled defaults, so the adopted snapshot is exactly what the files say. From f129d5775f2ba700c4997d37e3999e5b7f062849 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:36:56 -0400 Subject: [PATCH 015/108] docs(config): no backend has a sub-block yet; index backends.go in AGENTS.md Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 16595721..a7e4a28e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,7 +34,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6564775a..0aac1259 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index ee264bf4..193a6c21 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -48,7 +48,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. ### Server From 58d9e77242f8cd30611d74b90285963dd76dc785 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:46:07 -0400 Subject: [PATCH 016/108] fix(dedupe)!: reserve, commit or release ids keyed by tenant and table Replace CheckAndMark with a two-phase Reserve -> Commit | Release contract with a lease on the pending claim, keyed by (tenant, table, id) under a versioned layout, and add dedupetest, the conformance suite every backend runs. - Pebble claims under a sharded in-memory lock, so concurrent requests with one id publish it once (#390). - Ingest reserves after encoding, publishes, then commits, releasing the id when the publish fails, so a retried 503 is published, not dropped (#384's loss; F2 closes the uncertain-publish window). An id held by another request answers 503 with the lease as Retry-After. - The same id in two tables is two ids (#222); an explicit null id is a missing id (#370). BREAKING: the key layout changes, so ids seen before the upgrade are accepted once more. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 1 + docs/src/content/docs/api.md | 6 +- docs/src/content/docs/architecture.md | 12 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 4 +- internal/api/ingest.go | 139 ++++++-- internal/api/ingest_test.go | 219 +++++++++++++ internal/app/app_test.go | 18 +- internal/dedupe/conformance_test.go | 54 ++++ internal/dedupe/dedupe.go | 85 ++++- internal/dedupe/dedupetest/dedupetest.go | 318 +++++++++++++++++++ internal/dedupe/embedded.go | 196 ++++++++++-- internal/dedupe/embedded_test.go | 59 +++- internal/dedupe/export_test.go | 39 +++ internal/dedupe/key.go | 71 +++++ internal/dedupe/managed.go | 126 +++++++- internal/dedupe/managed_test.go | 97 ++++-- internal/dedupe/stores_test.go | 18 +- internal/settings/validate.go | 7 +- internal/settings/validate_test.go | 1 + internal/testutil/mocks.go | 99 +++++- 22 files changed, 1424 insertions(+), 151 deletions(-) create mode 100644 internal/dedupe/conformance_test.go create mode 100644 internal/dedupe/dedupetest/dedupetest.go create mode 100644 internal/dedupe/export_test.go create mode 100644 internal/dedupe/key.go diff --git a/AGENTS.md b/AGENTS.md index 16595721..47938eaf 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` @@ -58,7 +58,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. 6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format` — a pre-v2 envelope carries none — or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. -8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. +8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase — `Reserve` → publish → `Commit`, or `Release` when the publish fails; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c715..038096e1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,6 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble` until a later sweep drops them. A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 1634aaab..9040a6bb 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -262,7 +262,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 400 | `{"error":"invalid json"}` | Malformed request body | | 400 | `{"error":"unknown column ... for table ..."}` (also: `missing required column ...`, `type mismatch for column ...`, `null value for non-nullable column ...`) | Schema validation failure (unknown fields, type mismatches, missing required columns, null in a non-nullable column with no default). The body is the validator's message verbatim — there is no `validation failed:` prefix. | | 400 | `{"error":"column \"x\" of table \"t\" is materialized and cannot be inserted"}` (also `… is alias …`) | The record supplies a value for a column ClickHouse computes. Omit it — the server fills it in. Refused rather than dropped: the published row has one slot per insertable column, so the value would otherwise vanish behind a `200` | -| 400 | `{"error":"missing dedupe id field \"event_id\""}` | Only when dedupe is enabled with `dedupe.require_id: true` and the row lacks the configured `id_field`. With `require_id: false` (the default) the row is instead published un-deduped. Either way — reject or publish — the row is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`. In a batch this is a per-record failure, not a whole-request error. | +| 400 | `{"error":"missing dedupe id field \"event_id\""}` | Only when dedupe is enabled with `dedupe.require_id: true` and the row lacks the configured `id_field` or sets it to `null`. With `require_id: false` (the default) the row is instead published un-deduped. Either way — reject or publish — the row is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`. In a batch this is a per-record failure, not a whole-request error. | | 401 | `{"error":"invalid token"}` / `{"error":"token expired"}` | A present-but-invalid/expired token was supplied and denied (the gate surfaces the token reason rather than silently falling back to `default_role`) | | 403 | `{"error":"forbidden"}` (empty-role variant: `forbidden: request has no role and no public default_role is configured`) | The resolved role lacks `insert` on the table | | 403 | `{"error":"column \"x\" not allowed for insert"}` | The record names a column the role's `allow_columns`/`deny_columns` forbids ([Access control → Column permissions](/access-control#column-permissions)) | @@ -274,7 +274,8 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error | -| 503 | `{"error":"service unavailable"}` | NATS JetStream stream full (backpressure). Response includes `Retry-After: 30` header. | +| 503 | `{"error":"service unavailable"}` | NATS JetStream stream full (backpressure). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** @@ -386,6 +387,7 @@ A `200` is returned whenever the body was read and the records were processed | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch | | 503 | `{"error":"service unavailable"}` | NATS JetStream full (backpressure) mid-batch; includes `Retry-After: 30` | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54..19d25bed 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -58,7 +58,7 @@ internal/ ├── chconn/ One ClickHouse pool per connection tuple among the served tenants, reconciled on reload under the ceiling ├── chsql/ Shared ClickHouse SQL helpers (identifier quoting, bind-safety) ├── config/ YAML + env var configuration loading -├── dedupe/ Optional deduplication (Pebble) +├── dedupe/ Optional deduplication (Reserve/Commit/Release; Pebble) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) @@ -80,7 +80,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. -- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). +- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup (the id reserved once the record is encoded, committed after the publish, released if the publish fails; an id another request holds answers `503` with the lease as `Retry-After`), and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. @@ -122,9 +122,11 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `dedupe/` — Deduplication (Optional) -- **dedupe.go** — `Deduplicator` interface: `CheckAndMark(ctx, eventID) (bool, error)`. -- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3), key = tenant id, a NUL, event id — no tenant id holds a NUL, so no two tenants' keys meet. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. -- **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight `CheckAndMark` calls are serialized against the swap, so flipping the key is a reload, not a restart. `CheckAndMark` returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). +- **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. +- **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. +- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes a window of claims in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key and collapses a key repeated in one call before the backend sees it, once for every backend. +- **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. ### `discovery/` — Schema Discovery & Validation diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 095b4990..63cff252 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -169,7 +169,7 @@ WH_SETTINGS_DIR=/etc/wavehouse/settings WaveHouse keeps all embedded state under a single configurable root, `WH_DATA_DIR` (yaml: `data_dir`). Subdirectories are convention, not config: - `/nats` — embedded NATS JetStream. Holds in-flight events between an ingest POST and the ingest worker → ClickHouse flush, plus the `stream.gap_window_minutes` window (settings directory) of history that powers SSE gap-fill across restarts. -- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). +- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). In a Docker / Podman / Kubernetes deployment, **`data_dir` must resolve to a host-backed volume**. The reference compose file `deployments/compose/standalone.yaml` sets `WH_DATA_DIR=/app/data` and binds a `wavehouse-data:/app/data` volume — copy that pattern. The bundled Dockerfiles pre-create `/app/data` and `/app/settings` owned by the nonroot user (UID 65532); the binary creates the `nats/` and `pebble/` subdirectories under `/app/data` itself on first run. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e7a19b..e8c31e15 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,8 +186,8 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. -- `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. +- `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. ## ClickHouse diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 029799b6..884638a0 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -7,8 +7,10 @@ import ( "fmt" "io" "log/slog" + "math" "net/http" "sort" + "strconv" "strings" "time" @@ -49,8 +51,11 @@ type IngestHandler struct { // record boundary — one record never mixes two documents' values. Dedup is // skipped when nil. DedupeSettings func(store *settings.Store, table string) (enabled bool, idField string, requireID bool) - Publisher mq.Publisher - PolicySource PolicySource + // DedupeLease is how long a record's claimed id stays pending while it is + // published; 0 means dedupe.DefaultLease. + DedupeLease time.Duration + Publisher mq.Publisher + PolicySource PolicySource // Validator and Checker are the per-record seams a native type layer will // take over (see ingest_seams.go). Both are optional: nil means the default @@ -74,6 +79,13 @@ var dedupeMissingIDCounter, _ = otel.Meter("wavehouse-ingest").Int64Counter( metric.WithDescription("Ingested records missing the configured dedupe id_field (idempotency skipped)"), ) +// dedupeCommitFailedCounter counts records published whose id could not be +// committed afterwards: a retry after the lease lapses publishes them again. +var dedupeCommitFailedCounter, _ = otel.Meter("wavehouse-ingest").Int64Counter( + "wavehouse_dedupe_commit_failed_total", + metric.WithDescription("Published records whose dedupe id failed to commit afterwards (the claim lapses with its lease)"), +) + // dedupeDisabledCounter counts records published un-deduped because the // settings snapshot said dedupe was on while the store was switched off — // transient across a reload; a climbing rate means the store and the @@ -125,7 +137,7 @@ type recordReject struct { // // Most causes are TRANSIENT system conditions, where abandoning the tail is what // makes the batch safe to retry: publish backpressure (503), a publish/marshal -// failure (500), a dedup backend error (500). +// failure (500), a dedup backend error (500), an id another request holds (503). // // One is not. An insert grant that resolved for the other operation is a 403 and // a caller/config bug — retrying cannot help. It aborts rather than rejecting @@ -639,41 +651,26 @@ func (h *IngestHandler) processRecord( // from one snapshot (table override → global; the settings directory // always states them, so no compiled fallback is needed), so a reload // lands at a record boundary. A Deduplicator without a settings source is - // a wiring bug, not a mode — main wires both or neither. + // a wiring bug, not a mode — main wires both or neither. The id is claimed + // only once the record is encoded, so nothing but the publish can fail + // while the claim is held. + var dedupKey *dedupe.Key if h.Dedup != nil && h.DedupeSettings != nil { if enabled, idField, requireID := h.DedupeSettings(store, table); enabled { - idVal, ok := data[idField] - if !ok { + // An explicit null is as missing as an absent key (#370): fmt.Sprint + // would make every null "", one id for every such record. + if idVal, ok := data[idField]; ok && idVal != nil { + dedupKey = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} + } else { dedupeMissingIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) if requireID { - slog.WarnContext(ctx, "dedupe id_field missing; rejecting", "id_field", idField, "table", table) + slog.WarnContext(ctx, "dedupe id_field missing or null; rejecting", "id_field", idField, "table", table) return false, &recordReject{ Status: http.StatusBadRequest, Message: fmt.Sprintf("missing dedupe id field %q", idField), }, nil } - slog.WarnContext(ctx, "dedupe id_field missing; publishing without idempotency", "id_field", idField, "table", table) - } else { - eventID := fmt.Sprint(idVal) - dup, err := h.Dedup(store).CheckAndMark(ctx, eventID) - switch { - case errors.Is(err, dedupe.ErrDisabled): - // A reload flipped dedupe.enabled between the snapshot - // read above and this call (the two transition at - // different instants). Publish un-deduped, as a record - // under the other setting would have been. The counter - // carries the signal (a burst is a reload; a steady rate - // is the store and settings out of step), so the line is - // Debug rather than a WARN per record. - dedupeDisabledCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) - slog.DebugContext(ctx, "dedupe switched off mid-reload; publishing without idempotency", "event_id", eventID, "table", table) - case err != nil: - slog.ErrorContext(ctx, "dedupe check failed", "error", err, "event_id", eventID) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} - case dup: - slog.InfoContext(ctx, "duplicate event skipped", "event_id", eventID) - return true, nil, nil - } + slog.WarnContext(ctx, "dedupe id_field missing or null; publishing without idempotency", "id_field", idField, "table", table) } } } @@ -704,8 +701,23 @@ func (h *IngestHandler) processRecord( return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} } + var dd dedupe.Deduplicator + var claims []dedupe.Claim + if dedupKey != nil { + dd = h.Dedup(store) + var duplicate bool + var abort *requestAbort + claims, duplicate, abort = h.reserve(ctx, dd, *dedupKey) + if duplicate || abort != nil { + return duplicate, nil, abort + } + } + slog.DebugContext(ctx, "publishing event to the ingest queue", "table", table, "scope", scope) if err := h.Publisher.Publish(ctx, mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope}, payload); err != nil { + // The record is not in the queue, so its id goes back: the client's + // retry must not read as a duplicate of it (#384). + releaseClaims(ctx, dd, claims) if errors.Is(err, mq.ErrQueueFull) { slog.WarnContext(ctx, "ingest queue is full", "error", err, "table", table, "scope", scope) return false, nil, &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} @@ -713,10 +725,77 @@ func (h *IngestHandler) processRecord( slog.ErrorContext(ctx, "failed to publish to the ingest queue", "error", err, "table", table, "scope", scope) return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "publish failed"} } - + commitClaims(ctx, dd, claims, table) return false, nil, nil } +// reserve claims key for one record. A duplicate skips the record; a key +// another request holds aborts with 503 and the lease as Retry-After, since +// that request's outcome decides this one's. ErrDisabled — a reload switched +// the store off after the settings snapshot was read — publishes un-deduped, +// as a record under the other setting would have been. +func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, key dedupe.Key) (claims []dedupe.Claim, duplicate bool, abort *requestAbort) { + lease := h.DedupeLease + if lease <= 0 { + lease = dedupe.DefaultLease + } + claims, err := dd.Reserve(ctx, []dedupe.Key{key}, lease) + switch { + case errors.Is(err, dedupe.ErrDisabled): + // The counter carries the signal (a burst is a reload; a steady rate + // is the store and settings out of step), so the line is Debug rather + // than a WARN per record. + dedupeDisabledCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", key.Table))) + slog.DebugContext(ctx, "dedupe switched off mid-reload; publishing without idempotency", "event_id", key.ID, "table", key.Table) + return nil, false, nil + case err != nil: + slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "event_id", key.ID, "table", key.Table) + return nil, false, &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} + } + switch claims[0].Status { + case dedupe.Duplicate: + slog.InfoContext(ctx, "duplicate event skipped", "event_id", key.ID, "table", key.Table) + return nil, true, nil + case dedupe.InFlight: + slog.InfoContext(ctx, "event id in flight in another request", "event_id", key.ID, "table", key.Table) + return nil, false, &requestAbort{ + Status: http.StatusServiceUnavailable, + Message: "a request with the same dedupe id is in flight", + RetryAfter: strconv.Itoa(int(math.Ceil(lease.Seconds()))), + } + case dedupe.Claimed: + } + return claims, false, nil +} + +// commitClaims makes a published record's id a duplicate. A failure does not +// fail the record — it is in the queue — so it is logged and counted, and +// the claim lapses after its lease. +func commitClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim, table string) { + if len(claims) == 0 { + return + } + // The record is queued whatever the request's context does next. + err := dd.Commit(context.WithoutCancel(ctx), claims, 0) + switch { + case err == nil, errors.Is(err, dedupe.ErrDisabled): + default: + dedupeCommitFailedCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) + slog.ErrorContext(ctx, "dedupe commit failed after publish; the id lapses with its lease", "error", err, "table", table) + } +} + +// releaseClaims gives claims back after a failed publish. A failure is only +// logged: the claim lapses with its lease either way. +func releaseClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim) { + if len(claims) == 0 { + return + } + if err := dd.Release(context.WithoutCancel(ctx), claims); err != nil && !errors.Is(err, dedupe.ErrDisabled) { + slog.WarnContext(ctx, "dedupe release failed; the id lapses with its lease", "error", err) + } +} + // checkValueMatches decides insert-check equality: the payload value must // have a canonical scalar form (object/array/null match nothing) equal to the // required value's canonical form. A policy.LiteralValue — and only that type, diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index 2ae205e3..f08d2eb0 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -11,6 +11,7 @@ import ( "net/http/httptest" "net/url" "strings" + "sync" "testing" "testing/iotest" "time" @@ -2737,3 +2738,221 @@ func TestIngest_CheckOnEphemeralColumn_Rejected(t *testing.T) { assert.Contains(t, jsonErrorMessage(t, w), "is ephemeral and is never stored") assert.Empty(t, pub.Messages, "an unenforceable check must publish nothing") } + +// dedupHandler is a handler over the clicks registry with dedupe on for +// event_id and dedup as the store. +func dedupHandler(t *testing.T, pub *testutil.MockPublisher, dedup dedupe.Deduplicator, requireID bool) *IngestHandler { + t.Helper() + h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) + h.Dedup = staticDedup(dedup) + h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", requireID } + return h +} + +// #384: a publish that fails gives the id back, so the retry the 503 asks for +// is published rather than skipped as a duplicate of a record that never +// reached the queue. +func TestIngest_Dedup_FailedPublishReleasesTheID(t *testing.T) { + t.Parallel() + tests := []struct { + name string + err error + status int + }{ + {"backpressure", fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull), http.StatusServiceUnavailable}, + {"other failure", errors.New("connection reset"), http.StatusInternalServerError}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: tt.err} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + body := map[string]any{"page": "/home", "event_id": "e1"} + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, tt.status, w.Code) + assert.False(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "released, not left to lapse") + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusOK, w.Code) + assert.Contains(t, w.Body.String(), `"ok":true`, "the retry is published, not a duplicate") + assert.Len(t, pub.Published(), 1) + + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + assert.Contains(t, w.Body.String(), `"duplicate":true`, "and committed once published") + }) + } +} + +// A batch whose publish fails part-way keeps what it published: the records +// before the failure are committed, the failing one is released, and a +// whole-batch retry reports the first as duplicates and publishes the rest. +func TestIngest_NDJSON_Dedup_PublishFailureMidBatch(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull), ErrAfter: 1} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + batch := func() *http.Request { + return ndjsonRequest(t, "clicks", + jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), + jsonLine(t, map[string]any{"page": "/b", "event_id": "e2"}), + jsonLine(t, map[string]any{"page": "/c", "event_id": "e3"}), + ) + } + + w := httptest.NewRecorder() + h.Handle(w, withTenant(batch())) + require.Equal(t, http.StatusServiceUnavailable, w.Code) + require.Len(t, dedup.Released, 1) + assert.Equal(t, dedupe.Key{Table: "clicks", ID: "e2"}, dedup.Released[0].Key) + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(batch())) + require.Equal(t, http.StatusOK, w.Code) + resp := decodeBatchResult(t, w) + assert.Equal(t, 1, resp.Duplicates, "e1 was published by the first attempt") + assert.Equal(t, 2, resp.Succeeded) + assert.Len(t, pub.Published(), 3, "every record exactly once") +} + +// An id another request holds answers 503 with the lease as Retry-After: +// that request's publish decides whether this record is a duplicate. +func TestIngest_Dedup_InFlight(t *testing.T) { + t.Parallel() + tests := []struct { + name string + lease time.Duration + retryAfter string + }{ + {"default lease", 0, "30"}, + {"configured lease rounds up", 4500 * time.Millisecond, "5"}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.Hold(dedupe.Key{Table: "clicks", ID: "e1"}) + h := dedupHandler(t, pub, dedup, false) + h.DedupeLease = tt.lease + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, tt.retryAfter, w.Header().Get("Retry-After")) + assert.Contains(t, w.Body.String(), "in flight") + assert.Empty(t, pub.Published()) + }) + } +} + +// A commit that fails after the publish does not fail the record: it is in +// the queue, and answering an error would invite a second copy. +func TestIngest_Dedup_CommitFailureStillSucceeds(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.CommitErr = errors.New("disk full") + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}))) + assert.Equal(t, http.StatusOK, w.Code) + assert.Len(t, pub.Published(), 1) +} + +// A dedupe backend error before the publish publishes nothing and fails the +// request, as before. +func TestIngest_Dedup_ReserveError(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.Err = errors.New("backend down") + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}))) + assert.Equal(t, http.StatusInternalServerError, w.Code) + assert.Contains(t, w.Body.String(), "dedupe failed") + assert.Empty(t, pub.Published()) +} + +// #370: an explicit null id is a missing id — rejected under require_id, +// published un-deduped otherwise — never the one id "" that made every +// null record after the first a duplicate. +func TestIngest_Dedup_NullIDIsMissing(t *testing.T) { + t.Parallel() + nullID := func() string { return jsonLine(t, map[string]any{"page": "/a", "event_id": nil}) } + t.Run("require_id rejects", func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, testutil.NewMockDeduplicator(), true) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", nullID()))) + require.Equal(t, http.StatusOK, w.Code) + assert.Contains(t, resultAt(t, decodeBatchResult(t, w), 1).Error, "missing dedupe id field") + assert.Empty(t, pub.Published()) + }) + t.Run("otherwise publishes every one", func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, testutil.NewMockDeduplicator(), false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", nullID(), nullID()))) + require.Equal(t, http.StatusOK, w.Code) + resp := decodeBatchResult(t, w) + assert.Equal(t, 2, resp.Succeeded) + assert.Equal(t, 0, resp.Duplicates) + assert.Len(t, pub.Published(), 2) + }) +} + +// #222: the key carries the table, so one id value in two tables is two ids. +func TestIngest_Dedup_KeyedByTable(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}))) + require.Equal(t, http.StatusOK, w.Code) + claims, err := dedup.Reserve(t.Context(), []dedupe.Key{{Table: "clicks", ID: "e1"}, {Table: "views", ID: "e1"}}, time.Second) + require.NoError(t, err) + assert.Equal(t, dedupe.Duplicate, claims[0].Status) + assert.Equal(t, dedupe.Claimed, claims[1].Status) +} + +// #390: concurrent requests carrying one id publish it once, over the real +// embedded store — the rest answer duplicate, or 503 while the winner is +// still publishing. +func TestIngest_Dedup_ConcurrentSameIDPublishesOnce(t *testing.T) { + t.Parallel() + store := dedupe.NewEmbedded(t.TempDir()).Tenant("acme") + require.NoError(t, store.Apply(true)) + t.Cleanup(func() { _ = store.Close() }) + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, store, false) + + const n = 32 + codes := make([]int, n) + var wg sync.WaitGroup + for i := range n { + wg.Go(func() { + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}))) + codes[i] = w.Code + }) + } + wg.Wait() + assert.Len(t, pub.Published(), 1) + for _, c := range codes { + assert.Contains(t, []int{http.StatusOK, http.StatusServiceUnavailable}, c) + } +} diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d9..37efd618 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -29,6 +29,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/cache" "github.com/Wave-RF/WaveHouse/internal/config" "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/dedupe/dedupetest" "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/settings" "github.com/Wave-RF/WaveHouse/internal/tenant" @@ -490,14 +491,14 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { for _, id := range []string{"acme", "globex", "broken"} { assert.NoDirExists(t, filepath.Join(cfg.DataDir, id), "and no directory of a tenant's own") } - dup, err := acme.CheckAndMark(ctx, "e1") + dup, err := dedupetest.Mark(ctx, acme, eventKey) require.NoError(t, err) assert.False(t, dup) rewriteSettings(t, filepath.Join(root, "globex"), dedupeOn) a.tenants.Reload("test") assert.True(t, globex.Open(), "globex's reload opens globex's store") - dup, err = globex.CheckAndMark(ctx, "e1") + dup, err = dedupetest.Mark(ctx, globex, eventKey) require.NoError(t, err) assert.False(t, dup, "an id acme has seen is new to globex") @@ -527,7 +528,7 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { a.tenants.Reload("test") restored := a.dedup.For("acme") assert.True(t, restored.Open()) - dup, err = restored.CheckAndMark(ctx, "e1") + dup, err = dedupetest.Mark(ctx, restored, eventKey) require.NoError(t, err) assert.True(t, dup, "an id seen before the folder was removed is still a duplicate") @@ -563,10 +564,10 @@ func TestNew_DedupeOpenFailure(t *testing.T) { for _, id := range []tenant.ID{"acme", "globex"} { store := a.dedup.For(id) assert.False(t, store.Open()) - _, err := store.CheckAndMark(t.Context(), "e1") + _, err := dedupetest.Mark(t.Context(), store, eventKey) require.ErrorIs(t, err, dedupe.ErrUnavailable, "%s: switched on but not open, so its ingest fails closed", id) } - _, err := a.dedup.For("initech").CheckAndMark(t.Context(), "e1") + _, err := dedupetest.Mark(t.Context(), a.dedup.For("initech"), eventKey) require.ErrorIs(t, err, dedupe.ErrDisabled, "a tenant with dedupe off is as it would be anyway") }) } @@ -1421,7 +1422,7 @@ func TestReload_TenantGoneReleasesItsPoolAndRegistry(t *testing.T) { a.Handler().ServeHTTP(rec, req) return fmt.Sprintf("%d %s", rec.Code, rec.Body.String()) } - dup, err := acmeDedup.CheckAndMark(t.Context(), "e1") + dup, err := dedupetest.Mark(t.Context(), acmeDedup, eventKey) require.NoError(t, err) require.False(t, dup) require.Eventually(t, func() bool { return fetches.Load() > 0 }, 5*time.Second, 10*time.Millisecond, "acme's key set is fetched off the boot path") @@ -1465,7 +1466,7 @@ func TestReload_TenantGoneReleasesItsPoolAndRegistry(t *testing.T) { assert.NotNil(t, a.discoveries.For("acme")) assert.NotSame(t, acmeRegistry, a.discoveries.For("acme"), "and a fresh registry") assert.Eventually(t, func() bool { return fetches.Load() > fetched }, 5*time.Second, 10*time.Millisecond, "and a fresh verifier, fetching the key set again") - dup, err = a.dedup.For("acme").CheckAndMark(t.Context(), "e1") + dup, err = dedupetest.Mark(t.Context(), a.dedup.For("acme"), eventKey) require.NoError(t, err) assert.True(t, dup, "an id acme sent before the removal is still a duplicate") } @@ -1557,3 +1558,6 @@ func TestClose_StopsTheDiscoveryLoops(t *testing.T) { assert.Nil(t, a.discoveries.For("acme")) assert.Nil(t, a.pools.For("acme")) } + +// eventKey is the one dedupe key the tenant-lifecycle tests mark. +var eventKey = dedupe.Key{Table: "events", ID: "e1"} diff --git a/internal/dedupe/conformance_test.go b/internal/dedupe/conformance_test.go new file mode 100644 index 00000000..afadae28 --- /dev/null +++ b/internal/dedupe/conformance_test.go @@ -0,0 +1,54 @@ +package dedupe_test + +import ( + "sync" + "testing" + "time" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/dedupe/dedupetest" +) + +// fakeClock is a clock tests move by hand. +type fakeClock struct { + mu sync.Mutex + now time.Time +} + +func (c *fakeClock) Now() time.Time { + c.mu.Lock() + defer c.mu.Unlock() + return c.now +} + +func (c *fakeClock) Advance(d time.Duration) { + c.mu.Lock() + defer c.mu.Unlock() + c.now = c.now.Add(d) +} + +func TestEmbedded_Conformance(t *testing.T) { + t.Parallel() + dedupetest.Run(t, func(t *testing.T) dedupetest.Harness { + e := dedupe.NewEmbedded(t.TempDir()) + clock := &fakeClock{now: time.Now()} + dedupe.SetClock(e, clock.Now) + return dedupetest.Harness{ + Factory: e.Tenant, + Advance: clock.Advance, + FailNextReserve: func(n int) { dedupe.FailNextReserve(e, n) }, + } + }) +} + +// The suite's sleeping path, which a backend without an injectable clock +// takes, on the real clock. +func TestEmbedded_ConformanceRealClock(t *testing.T) { + t.Parallel() + if testing.Short() { + t.Skip("sleeps past leases") + } + dedupetest.Run(t, func(t *testing.T) dedupetest.Harness { + return dedupetest.Harness{Factory: dedupe.NewEmbedded(t.TempDir()).Tenant} + }) +} diff --git a/internal/dedupe/dedupe.go b/internal/dedupe/dedupe.go index c9fdb7f4..7602443a 100644 --- a/internal/dedupe/dedupe.go +++ b/internal/dedupe/dedupe.go @@ -1,13 +1,84 @@ package dedupe -import "context" +import ( + "context" + "time" +) -// Deduplicator checks whether an event has been seen before and marks it. -type Deduplicator interface { - // CheckAndMark returns true if the event was already seen (duplicate). - // If not seen, it atomically marks the event as seen. - CheckAndMark(ctx context.Context, eventID string) (isDuplicate bool, err error) +// DefaultLease is how long a Claimed key stays pending when the caller names +// no lease: long enough to cover a publish, short enough that a request that +// died mid-publish does not hold the id for long. +const DefaultLease = 30 * time.Second + +// Key is one record's dedupe identity inside a tenant's store. The tenant is +// bound by the store (Stores.For), so a Key never carries it. +type Key struct { + Table string + ID string +} + +// Status is Reserve's verdict for one key. +type Status uint8 + +const ( + // Claimed is a first sighting within retention. The caller now holds a + // pending claim and must Commit it once the record is published, or + // Release it if the publish definitely failed. An abandoned claim lapses + // after the lease. + Claimed Status = iota + 1 + // Duplicate means the key was committed earlier and has not expired: skip + // the record. Also returned for a key repeated inside one Reserve call, + // after its first occurrence. + Duplicate + // InFlight means another request holds a live claim on the key. Its + // outcome is not known yet, so the caller answers 503 and the client + // retries. + InFlight +) - // Close releases resources held by the deduplicator. +func (s Status) String() string { + switch s { + case Claimed: + return "claimed" + case Duplicate: + return "duplicate" + case InFlight: + return "in_flight" + default: + return "unknown" + } +} + +// Claim is Reserve's answer for one key. Token is the backend's proof of +// ownership, opaque to callers; Release compares it. +type Claim struct { + Key Key + Status Status + Token string +} + +// Deduplicator is a tenant's store of seen ids. +// +// Reserve is atomic per key: of any number of concurrent Reserves for the +// same key — in this process or any other sharing the backend — at most one +// returns Claimed. It returns one Claim per key, in input order. On error it +// has released every claim it made (all-or-nothing from the caller's view), +// and the error wraps ErrUnavailable when retrying later can succeed +// (throttled, timed out, backend unreachable). +// +// Commit makes Claimed claims duplicates for retention (0 = no expiry) and +// ignores claims of any other status. It is unconditional: a commit that +// lands after its lease lapsed and another request re-claimed the key is +// still correct, because the committing request did publish. +// +// Release gives up the Claimed claims it still owns (token match); a claim +// that has lapsed or been re-claimed is left alone. +// +// There is deliberately no read-only check: every caller that asks "have I +// seen this" needs the claim too, and a separate read is how #390 happened. +type Deduplicator interface { + Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) + Commit(ctx context.Context, claims []Claim, retention time.Duration) error + Release(ctx context.Context, claims []Claim) error Close() error } diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go new file mode 100644 index 00000000..c3d1cd9c --- /dev/null +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -0,0 +1,318 @@ +// Package dedupetest is the conformance suite every dedupe backend runs: the +// Deduplicator contract (dedupe.go) as tests, driven through the production +// path — a backend's Factory and the Managed switch it returns — so a backend +// that passes here behaves the same under ingest as every other. +package dedupetest + +import ( + "context" + "errors" + "fmt" + "strings" + "sync" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// Harness is one fresh backend under test. +type Harness struct { + // Factory builds a tenant's store over the backend. The suite switches + // each store it builds on and closes it at cleanup. + Factory dedupe.Factory + // Peer, if set, builds a tenant's store over the same data through a + // second client — another process's view, for backends that have one. + // nil uses Factory. + Peer dedupe.Factory + // Advance moves the backend's clock forward by d. nil makes the suite + // sleep instead, which is why its leases and retentions are whole + // seconds: a backend may store expiry at one-second resolution. + Advance func(d time.Duration) + // FailNextReserve, if set, makes the backend's next Reserve fail after it + // has claimed n keys. nil skips the case that needs it. + FailNextReserve func(n int) +} + +// Run runs every case, each against a backend newHarness builds fresh. +func Run(t *testing.T, newHarness func(t *testing.T) Harness) { + t.Helper() + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + t.Parallel() + c.run(t, &suite{Harness: newHarness(t)}) + }) + } +} + +// Mark is the old check-and-mark in one call, for tests that only need an id +// seen: it reserves k and commits it with no expiry, reporting whether k was +// already committed. A key another request holds is an error. +func Mark(ctx context.Context, d dedupe.Deduplicator, k dedupe.Key) (duplicate bool, err error) { + claims, err := d.Reserve(ctx, []dedupe.Key{k}, dedupe.DefaultLease) + if err != nil { + return false, err + } + switch claims[0].Status { + case dedupe.Duplicate: + return true, nil + case dedupe.Claimed: + return false, d.Commit(ctx, claims, 0) + case dedupe.InFlight: + } + return false, fmt.Errorf("key %v is %s", k, claims[0].Status) +} + +const ( + lease = time.Second + // long outlives every case, so only a deliberate pass lapses it. + long = time.Hour +) + +type suite struct { + Harness +} + +func (s *suite) open(t *testing.T, build dedupe.Factory, id tenant.ID) dedupe.Deduplicator { + t.Helper() + m := build(id) + require.NoError(t, m.Apply(true)) + t.Cleanup(func() { _ = m.Close() }) + return m +} + +// store opens tenant id's store; peer opens it through the second client. +func (s *suite) store(t *testing.T, id tenant.ID) dedupe.Deduplicator { + t.Helper() + return s.open(t, s.Factory, id) +} + +func (s *suite) peer(t *testing.T, id tenant.ID) dedupe.Deduplicator { + t.Helper() + if s.Peer == nil { + return s.store(t, id) + } + return s.open(t, s.Peer, id) +} + +// pass lets d go by, plus a second's margin for a backend that stores expiry +// in whole seconds. +func (s *suite) pass(d time.Duration) { + d += time.Second + if s.Advance != nil { + s.Advance(d) + return + } + time.Sleep(d) +} + +func reserve(t *testing.T, d dedupe.Deduplicator, lease time.Duration, keys ...dedupe.Key) []dedupe.Claim { + t.Helper() + claims, err := d.Reserve(t.Context(), keys, lease) + require.NoError(t, err) + require.Len(t, claims, len(keys)) + for i, c := range claims { + require.Equal(t, keys[i], c.Key, "claim %d answers its own key, in input order", i) + } + return claims +} + +func statuses(claims []dedupe.Claim) []dedupe.Status { + out := make([]dedupe.Status, len(claims)) + for i, c := range claims { + out[i] = c.Status + } + return out +} + +func key(id string) dedupe.Key { return dedupe.Key{Table: "events", ID: id} } + +var cases = []struct { + name string + run func(t *testing.T, s *suite) +}{ + {"claim then commit is a duplicate", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + c := reserve(t, d, long, key("e1")) + require.Equal(t, dedupe.Claimed, c[0].Status) + assert.NotEmpty(t, c[0].Token) + require.NoError(t, d.Commit(t.Context(), c, 0)) + assert.Equal(t, dedupe.Duplicate, reserve(t, d, long, key("e1"))[0].Status) + assert.Equal(t, dedupe.Claimed, reserve(t, d, long, key("e2"))[0].Status, "distinct ids are independent") + }}, + {"a released claim can be claimed again", func(t *testing.T, s *suite) { + // #384: a publish that failed releases the id, and the client's retry + // goes through. + d := s.store(t, "acme") + c := reserve(t, d, long, key("e1")) + require.NoError(t, d.Release(t.Context(), c)) + c = reserve(t, d, long, key("e1")) + assert.Equal(t, dedupe.Claimed, c[0].Status) + require.NoError(t, d.Commit(t.Context(), c, 0)) + assert.Equal(t, dedupe.Duplicate, reserve(t, d, long, key("e1"))[0].Status) + }}, + {"a live claim is in flight to everyone else", func(t *testing.T, s *suite) { + d, p := s.store(t, "acme"), s.peer(t, "acme") + c := reserve(t, d, long, key("e1")) + assert.Equal(t, dedupe.InFlight, reserve(t, d, long, key("e1"))[0].Status) + assert.Equal(t, dedupe.InFlight, reserve(t, p, long, key("e1"))[0].Status, "and to another client") + require.NoError(t, d.Commit(t.Context(), c, 0)) + assert.Equal(t, dedupe.Duplicate, reserve(t, p, long, key("e1"))[0].Status, "the peer sees the commit") + }}, + {"an abandoned claim lapses after its lease", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + reserve(t, d, lease, key("e1")) + s.pass(lease) + assert.Equal(t, dedupe.Claimed, reserve(t, d, long, key("e1"))[0].Status) + }}, + {"a commit expires after its retention", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + c := reserve(t, d, long, key("brief"), key("kept")) + require.NoError(t, d.Commit(t.Context(), c[:1], time.Second)) + require.NoError(t, d.Commit(t.Context(), c[1:], 0)) + assert.Equal(t, []dedupe.Status{dedupe.Duplicate, dedupe.Duplicate}, statuses(reserve(t, d, long, key("brief"), key("kept")))) + s.pass(time.Second) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Duplicate}, statuses(reserve(t, d, long, key("brief"), key("kept"))), + "retention 0 never expires") + }}, + {"concurrent reserves of one key claim it once", func(t *testing.T, s *suite) { + // #390: two requests carrying one id must not both publish. + const n = 64 + d, p := s.store(t, "acme"), s.peer(t, "acme") + race := func() []dedupe.Claim { + out := make([]dedupe.Claim, n) + var wg sync.WaitGroup + for i := range n { + store := d + if i%2 == 1 { + store = p + } + wg.Go(func() { + c, err := store.Reserve(context.Background(), []dedupe.Key{key("e1")}, long) + if assert.NoError(t, err) { + out[i] = c[0] + } + }) + } + wg.Wait() + return out + } + var winner []dedupe.Claim + for _, c := range race() { + if c.Status == dedupe.Claimed { + winner = append(winner, c) + } else { + assert.Equal(t, dedupe.InFlight, c.Status) + } + } + require.Len(t, winner, 1, "exactly one reserve claims the key") + require.NoError(t, d.Commit(t.Context(), winner, 0)) + for _, c := range race() { + assert.Equal(t, dedupe.Duplicate, c.Status) + } + }}, + {"a key repeated in one call is claimed once", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + c := reserve(t, d, long, key("a"), key("b"), key("a"), key("a")) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed, dedupe.Duplicate, dedupe.Duplicate}, statuses(c)) + require.NoError(t, d.Commit(t.Context(), c, 0), "commit ignores the repeats") + assert.Equal(t, []dedupe.Status{dedupe.Duplicate, dedupe.Duplicate}, statuses(reserve(t, d, long, key("a"), key("b")))) + }}, + {"answers keep input order in a large call", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + keys := make([]dedupe.Key, 300) + for i := range keys { + keys[i] = key(fmt.Sprint(i)) + } + var odd []dedupe.Key + for i := 1; i < len(keys); i += 2 { + odd = append(odd, keys[i]) + } + require.NoError(t, d.Commit(t.Context(), reserve(t, d, long, odd...), 0)) + for i, c := range reserve(t, d, long, keys...) { + want := dedupe.Claimed + if i%2 == 1 { + want = dedupe.Duplicate + } + assert.Equal(t, want, c.Status, "key %d", i) + } + }}, + {"tables and tenants have their own keyspace", func(t *testing.T, s *suite) { + acme, globex := s.store(t, "acme"), s.store(t, "globex") + // "ab"+"c" and "a"+"bc" would be one key were table and id just + // joined; tenants "a"/"ab" likewise. + first := []dedupe.Key{{Table: "clicks", ID: "e1"}, {Table: "ab", ID: "c"}} + require.NoError(t, acme.Commit(t.Context(), reserve(t, acme, long, first...), 0)) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed}, + statuses(reserve(t, acme, long, dedupe.Key{Table: "views", ID: "e1"}, dedupe.Key{Table: "a", ID: "bc"})), "#222: another table's id") + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed}, + statuses(reserve(t, globex, long, first...)), "another tenant's ids") + a, ab := s.store(t, "a"), s.store(t, "ab") + require.NoError(t, a.Commit(t.Context(), reserve(t, a, long, dedupe.Key{Table: "bt", ID: "e1"}), 0)) + assert.Equal(t, dedupe.Claimed, reserve(t, ab, long, dedupe.Key{Table: "t", ID: "e1"})[0].Status) + assert.Equal(t, []dedupe.Status{dedupe.Duplicate, dedupe.Duplicate}, + statuses(reserve(t, acme, long, first...)), "and still duplicates in their own") + }}, + {"long ids and ids that look hashed stay distinct", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + base := strings.Repeat("x", 2*dedupe.MaxIDBytes) + longA, longB := key(base+"a"), key(base+"b") + hashLike := key("\xff" + strings.Repeat("0", 32)) + require.NoError(t, d.Commit(t.Context(), reserve(t, d, long, longA, hashLike), 0)) + assert.Equal(t, []dedupe.Status{dedupe.Duplicate, dedupe.Claimed, dedupe.Duplicate}, + statuses(reserve(t, d, long, longA, longB, hashLike))) + }}, + {"a late commit after a re-claim still lands", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + first := reserve(t, d, lease, key("e1")) + s.pass(lease) + second := reserve(t, d, long, key("e1")) + require.Equal(t, dedupe.Claimed, second[0].Status) + require.NoError(t, d.Commit(t.Context(), first, 0), "the first request did publish") + require.NoError(t, d.Release(t.Context(), second), "the second gives up; the commit stands") + assert.Equal(t, dedupe.Duplicate, reserve(t, d, long, key("e1"))[0].Status) + }}, + {"a stale release leaves the new claimant alone", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + first := reserve(t, d, lease, key("e1")) + s.pass(lease) + second := reserve(t, d, long, key("e1")) + require.NoError(t, d.Release(t.Context(), first)) + assert.Equal(t, dedupe.InFlight, reserve(t, d, long, key("e1"))[0].Status, "the second claim is still live") + require.NoError(t, d.Commit(t.Context(), second, 0)) + assert.Equal(t, dedupe.Duplicate, reserve(t, d, long, key("e1"))[0].Status) + }}, + {"a failed reserve leaves nothing claimed", func(t *testing.T, s *suite) { + if s.FailNextReserve == nil { + t.Skip("the backend has no failure hook") + } + d := s.store(t, "acme") + s.FailNextReserve(1) + _, err := d.Reserve(t.Context(), []dedupe.Key{key("a"), key("b"), key("c")}, long) + require.Error(t, err) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed, dedupe.Claimed}, + statuses(reserve(t, d, long, key("a"), key("b"), key("c")))) + }}, + {"empty calls are no-ops", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + c, err := d.Reserve(t.Context(), nil, long) + require.NoError(t, err) + assert.Empty(t, c) + require.NoError(t, d.Commit(t.Context(), nil, 0)) + require.NoError(t, d.Release(t.Context(), nil)) + dup := []dedupe.Claim{{Key: key("e1"), Status: dedupe.Duplicate}, {Key: key("e2"), Status: dedupe.InFlight}} + require.NoError(t, d.Commit(t.Context(), dup, 0), "only Claimed claims commit") + assert.Equal(t, dedupe.Claimed, reserve(t, d, long, key("e1"))[0].Status) + }}, + {"a table holding NUL is refused", func(t *testing.T, s *suite) { + d := s.store(t, "acme") + _, err := d.Reserve(t.Context(), []dedupe.Key{key("ok"), {Table: "a\x00b", ID: "e1"}}, long) + require.ErrorIs(t, err, dedupe.ErrInvalidKey) + assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a bad key is not worth retrying") + assert.Equal(t, dedupe.Claimed, reserve(t, d, long, key("ok"))[0].Status, "and nothing was claimed") + }}, +} diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 6b3b5efa..b3f015f0 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -4,9 +4,13 @@ import ( "context" "encoding/binary" "errors" + "fmt" + "hash/fnv" "math" "path/filepath" + "strconv" "sync" + "sync/atomic" "time" "github.com/cockroachdb/pebble" @@ -16,23 +20,32 @@ import ( // Embedded is the embedded implementation: every tenant's seen ids in one // Pebble instance at data_dir/pebble, each key led by its tenant (#583 story -// 3), so a thousand tenants cost one instance's goroutines, open files and -// heap rather than a thousand. The instance opens with the first tenant's -// store switched on and closes with the last one switched off: it is open -// exactly while some tenant has dedupe on, and a tenant switched off, -// rejected or removed keeps its seen ids for when it is back. +// 3) and then its table (#222), with the pending claims in memory beside it +// (pendingSet). One instance means a thousand tenants cost one instance's +// goroutines, open files and heap rather than a thousand. The instance opens +// with the first tenant's store switched on and closes with the last one +// switched off: it is open exactly while some tenant has dedupe on, and a +// tenant switched off, rejected or removed keeps its seen ids for when it is +// back. Pebble is one process's, so two pods on it do not share seen ids. type Embedded struct { dir string mu sync.Mutex // guards db and open db *pebble.DB open int // tenant stores open over db + + pending *pendingSet + tokens atomic.Uint64 + now func() time.Time + // readHook, when set, runs before each Pebble read in Reserve; a test + // makes it fail to exercise Reserve's all-or-nothing error path. + readHook func() error } // NewEmbedded returns the embedded implementation under dataDir. Nothing is // opened until a tenant's store is. func NewEmbedded(dataDir string) *Embedded { - return &Embedded{dir: filepath.Join(dataDir, "pebble")} + return &Embedded{dir: filepath.Join(dataDir, "pebble"), pending: newPendingSet(), now: time.Now} } // Dir is where the instance lives. @@ -45,14 +58,10 @@ func (e *Embedded) Open() bool { return e.db != nil } -// keySeparator ends the tenant at the front of every key. A tenant id has no -// NUL, so the first one in a key is this one, and no two tenants' keys meet. -const keySeparator = 0 - // Tenant builds tenant id's store, closed, over its share of the instance — // the Factory Stores takes. func (e *Embedded) Tenant(id tenant.ID) *Managed { - prefix := append([]byte(id), keySeparator) + prefix := KeyPrefix(id) return NewManaged(func() (Deduplicator, error) { return e.acquire(prefix) }) } @@ -115,28 +124,115 @@ type tenantStore struct { closed sync.Once } -// CheckAndMark returns true if the event was already seen. -func (s *tenantStore) CheckAndMark(_ context.Context, eventID string) (bool, error) { - key := make([]byte, 0, len(s.prefix)+len(eventID)) - key = append(append(key, s.prefix...), eventID...) +// Committed values are committedMark ‖ expiry (big-endian UnixNano, 0 = +// never). Version-0 values were a bare 8-byte timestamp under version-0 +// keys, which no version-1 key reads. +const ( + committedMark = 2 + valueLen = 9 +) - _, closer, err := s.db.Get(key) - if err == nil { +// Reserve claims each key under its shard's lock: the pending check, the +// Pebble read and the claim happen with no other Reserve for that key in +// between, and Pebble's directory lock keeps a second process off the +// instance, so at most one caller holds a key (#390). +func (s *tenantStore) Reserve(_ context.Context, keys []Key, lease time.Duration) ([]Claim, error) { + now := s.e.now() + claims := make([]Claim, 0, len(keys)) + for _, k := range keys { + c, err := s.reserve(AppendKey(nil, s.prefix, k), k, now, lease) + if err != nil { + s.release(claims) + return nil, err + } + claims = append(claims, c) + } + return claims, nil +} + +func (s *tenantStore) reserve(key []byte, k Key, now time.Time, lease time.Duration) (Claim, error) { + sh := s.e.pending.shard(key) + sh.mu.Lock() + defer sh.mu.Unlock() + sh.sweep(now) + if p, ok := sh.m[string(key)]; ok && now.Before(p.expires) { + return Claim{Key: k, Status: InFlight}, nil + } + if s.e.readHook != nil { + if err := s.e.readHook(); err != nil { + return Claim{}, err + } + } + val, closer, err := s.db.Get(key) + switch { + case err == nil: + live := committedLive(val, now) _ = closer.Close() - return true, nil + if live { + return Claim{Key: k, Status: Duplicate}, nil + } + case !errors.Is(err, pebble.ErrNotFound): + return Claim{}, fmt.Errorf("dedupe read: %w", err) } - if !errors.Is(err, pebble.ErrNotFound) { - return false, err + token := strconv.FormatUint(s.e.tokens.Add(1), 36) + sh.m[string(key)] = pending{token: token, expires: now.Add(lease)} + return Claim{Key: k, Status: Claimed, Token: token}, nil +} + +// committedLive reports whether a stored value is a commit that has not +// expired. +func committedLive(val []byte, now time.Time) bool { + if len(val) != valueLen || val[0] != committedMark { + return false } + exp := int64(binary.BigEndian.Uint64(val[1:])) //nolint:gosec // written from an int64 below + return exp == 0 || now.UnixNano() < exp +} - // Store timestamp as value for future auditing. - val := make([]byte, 8) - binary.BigEndian.PutUint64(val, uint64(time.Now().UnixNano())) +// Commit writes every claim in one batch and one fsync, then drops the +// pending entries it still owns — in that order, so no Reserve in between +// finds the key neither pending nor committed. +func (s *tenantStore) Commit(_ context.Context, claims []Claim, retention time.Duration) error { + var exp int64 + if retention > 0 { + exp = s.e.now().Add(retention).UnixNano() + } + val := make([]byte, valueLen) + val[0] = committedMark + binary.BigEndian.PutUint64(val[1:], uint64(exp)) + b := s.db.NewBatch() + defer func() { _ = b.Close() }() + for _, c := range claims { + if err := b.Set(AppendKey(nil, s.prefix, c.Key), val, nil); err != nil { + return fmt.Errorf("dedupe commit: %w", err) + } + } + if err := b.Commit(pebble.Sync); err != nil { + return fmt.Errorf("dedupe commit: %w", err) + } + s.release(claims) + return nil +} - if err := s.db.Set(key, val, pebble.Sync); err != nil { - return false, err +// Release drops the pending entries the claims still own. +func (s *tenantStore) Release(_ context.Context, claims []Claim) error { + s.release(claims) + return nil +} + +func (s *tenantStore) release(claims []Claim) { + for _, c := range claims { + if c.Status != Claimed { + continue + } + key := AppendKey(nil, s.prefix, c.Key) + sh := s.e.pending.shard(key) + sh.mu.Lock() + if p, ok := sh.m[string(key)]; ok && p.token == c.Token { + delete(sh.m, string(key)) + } + sh.mu.Unlock() } - return false, nil } // Close releases the store's hold on the instance. Safe to call more than @@ -146,3 +242,51 @@ func (s *tenantStore) Close() error { s.closed.Do(func() { err = s.e.release() }) return err } + +// pendingShards spreads the pending claims over independently locked maps, +// so Reserves for different keys rarely wait on each other. +const pendingShards = 64 + +type pending struct { + token string + expires time.Time +} + +type pendingShard struct { + mu sync.Mutex + m map[string]pending + nextSweep time.Time +} + +// sweep drops lapsed claims at most once a DefaultLease, so a claim nobody +// commits, releases or re-reserves does not stay in memory. Callers hold mu. +func (sh *pendingShard) sweep(now time.Time) { + if now.Before(sh.nextSweep) { + return + } + sh.nextSweep = now.Add(DefaultLease) + for k, p := range sh.m { + if !now.Before(p.expires) { + delete(sh.m, k) + } + } +} + +// pendingSet is every tenant's live claims. It lives in memory because one +// process owns the instance: a crash forgets every claim, which is each +// lease lapsing at once. +type pendingSet [pendingShards]pendingShard + +func newPendingSet() *pendingSet { + p := new(pendingSet) + for i := range p { + p[i].m = map[string]pending{} + } + return p +} + +func (p *pendingSet) shard(key []byte) *pendingShard { + h := fnv.New32a() + _, _ = h.Write(key) + return &p[h.Sum32()%pendingShards] +} diff --git a/internal/dedupe/embedded_test.go b/internal/dedupe/embedded_test.go index de65a151..5015318f 100644 --- a/internal/dedupe/embedded_test.go +++ b/internal/dedupe/embedded_test.go @@ -2,8 +2,13 @@ package dedupe import ( "context" + "maps" "os" + "slices" "testing" + "time" + + "github.com/cockroachdb/pebble" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -26,15 +31,15 @@ func TestEmbedded_FirstSeenThenDuplicate(t *testing.T) { m := switchedOn(t, NewEmbedded(t.TempDir()), "acme") ctx := context.Background() - dup, err := m.CheckAndMark(ctx, "event-1") + dup, err := mark(ctx, m, "event-1") require.NoError(t, err) assert.False(t, dup, "first occurrence must not be a duplicate") - dup, err = m.CheckAndMark(ctx, "event-1") + dup, err = mark(ctx, m, "event-1") require.NoError(t, err) assert.True(t, dup, "second occurrence of the same id must be a duplicate") - dup, err = m.CheckAndMark(ctx, "event-2") + dup, err = mark(ctx, m, "event-2") require.NoError(t, err) assert.False(t, dup, "distinct ids are independent") } @@ -49,16 +54,16 @@ func TestEmbedded_TenantsDoNotShareSeenIDs(t *testing.T) { ctx := context.Background() a, ab := switchedOn(t, e, "a"), switchedOn(t, e, "ab") - dup, err := a.CheckAndMark(ctx, "bc") + dup, err := mark(ctx, a, "bc") require.NoError(t, err) assert.False(t, dup) - dup, err = ab.CheckAndMark(ctx, "c") + dup, err = mark(ctx, ab, "c") require.NoError(t, err) assert.False(t, dup, "another tenant's key, however the two would join") - dup, err = ab.CheckAndMark(ctx, "bc") + dup, err = mark(ctx, ab, "bc") require.NoError(t, err) assert.False(t, dup, "an id tenant a has seen is new to tenant ab") - dup, err = a.CheckAndMark(ctx, "bc") + dup, err = mark(ctx, a, "bc") require.NoError(t, err) assert.True(t, dup, "and still a duplicate within its own tenant") } @@ -77,7 +82,7 @@ func TestEmbedded_OpenWhileAnyTenantStoreIs(t *testing.T) { require.NoError(t, acme.Apply(true)) require.NoError(t, globex.Apply(true)) assert.True(t, e.Open()) - _, err := acme.CheckAndMark(ctx, "e1") + _, err := mark(ctx, acme, "e1") require.NoError(t, err) require.NoError(t, acme.Apply(false)) @@ -90,7 +95,7 @@ func TestEmbedded_OpenWhileAnyTenantStoreIs(t *testing.T) { require.NoError(t, acme.Apply(true)) t.Cleanup(func() { _ = acme.Close() }) - dup, err := acme.CheckAndMark(ctx, "e1") + dup, err := mark(ctx, acme, "e1") require.NoError(t, err) assert.True(t, dup, "a tenant switched off keeps its seen ids") } @@ -104,7 +109,7 @@ func TestEmbedded_StatsAreTheInstances(t *testing.T) { acme := switchedOn(t, e, "acme") switchedOn(t, e, "globex") - _, err := acme.CheckAndMark(context.Background(), "e1") + _, err := mark(context.Background(), acme, "e1") require.NoError(t, err) stats := e.Stats() m := e.db.Metrics() @@ -126,7 +131,7 @@ func TestEmbedded_OpenFailure(t *testing.T) { require.Error(t, acme.Apply(true)) require.Error(t, globex.Apply(true), "one instance: its failure is every tenant's") assert.False(t, e.Open()) - _, err := acme.CheckAndMark(context.Background(), "e1") + _, err := mark(context.Background(), acme, "e1") require.ErrorIs(t, err, ErrUnavailable) require.NoError(t, os.Remove(e.Dir())) @@ -134,3 +139,35 @@ func TestEmbedded_OpenFailure(t *testing.T) { t.Cleanup(func() { _ = acme.Close() }) assert.True(t, e.Open()) } + +// Keys from before the table joined the key (#222) are never read: an id +// seen then is accepted once more after the upgrade, the documented cost of +// the new layout. +func TestEmbedded_VersionZeroKeysAreNotRead(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + m := switchedOn(t, e, "acme") + require.NoError(t, e.db.Set([]byte("acme\x00e1"), make([]byte, 8), pebble.Sync)) + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.False(t, dup) +} + +// A claim nobody commits, releases or reserves again leaves memory at the +// next sweep past its lease, not never. +func TestPendingShard_SweepDropsLapsedClaims(t *testing.T) { + t.Parallel() + now := time.Now() + sh := &pendingShard{m: map[string]pending{ + "lapsed": {token: "1", expires: now.Add(-time.Second)}, + "live": {token: "2", expires: now.Add(time.Hour)}, + }} + sh.sweep(now) + assert.Equal(t, []string{"live"}, slices.Collect(maps.Keys(sh.m))) + + sh.m["lapsed"] = pending{token: "3", expires: now.Add(-time.Second)} + sh.sweep(now.Add(time.Second)) + assert.Len(t, sh.m, 2, "at most one sweep per DefaultLease") + sh.sweep(now.Add(DefaultLease)) + assert.Len(t, sh.m, 1) +} diff --git a/internal/dedupe/export_test.go b/internal/dedupe/export_test.go new file mode 100644 index 00000000..76e9dbc3 --- /dev/null +++ b/internal/dedupe/export_test.go @@ -0,0 +1,39 @@ +package dedupe + +import ( + "context" + "errors" + "sync/atomic" + "time" +) + +// SetClock replaces e's clock, for tests that let leases and retentions lapse +// without sleeping. +func SetClock(e *Embedded, now func() time.Time) { e.now = now } + +// FailNextReserve makes e's next Reserve fail after it has claimed n keys, +// once. +func FailNextReserve(e *Embedded, n int) { + var reads atomic.Int64 + var failed atomic.Bool + e.readHook = func() error { + if reads.Add(1) > int64(n) && failed.CompareAndSwap(false, true) { + return errors.New("injected read failure") + } + return nil + } +} + +// mark reserves and commits id in table "events", reporting whether it was +// already committed — the old check-and-mark, for tests about everything +// else. +func mark(ctx context.Context, d Deduplicator, id string) (bool, error) { + claims, err := d.Reserve(ctx, []Key{{Table: "events", ID: id}}, DefaultLease) + if err != nil { + return false, err + } + if claims[0].Status != Claimed { + return true, nil + } + return false, d.Commit(ctx, claims, 0) +} diff --git a/internal/dedupe/key.go b/internal/dedupe/key.go new file mode 100644 index 00000000..25ead7c2 --- /dev/null +++ b/internal/dedupe/key.go @@ -0,0 +1,71 @@ +package dedupe + +import ( + "crypto/sha256" + "errors" + "fmt" + "strings" + + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// The key layout every backend stores, byte for byte: +// +// keyVersion ‖ tenant ‖ keySeparator ‖ table ‖ keySeparator ‖ id +// +// A tenant id is letters, digits, '_' and '-', so it never holds the +// separator and never starts with keyVersion — the version-0 keys before +// #222 (tenant ‖ 0x00 ‖ id) never meet these. The id is last, so it may hold +// anything. +const ( + keyVersion byte = 0x01 + keySeparator byte = 0x00 + // hashedID leads an id stored as its SHA-256 rather than verbatim. Ids + // that start with it are hashed too, so a verbatim id never reads as a + // hashed one. + hashedID byte = 0xFF + // MaxIDBytes is the longest id stored verbatim: a DynamoDB partition key + // holds at most 2,048 bytes, and the tenant and table share them. + MaxIDBytes = 1024 +) + +// ErrInvalidKey is returned for a key no backend can store: a table name +// holding the separator byte. +var ErrInvalidKey = errors.New("invalid dedupe key") + +// KeyPrefix is the part of every key that names tenant id, so a backend +// computes it once per tenant store. +func KeyPrefix(id tenant.ID) []byte { + p := make([]byte, 0, len(id)+2) + p = append(p, keyVersion) + p = append(p, id...) + return append(p, keySeparator) +} + +// Validate reports whether k can be stored. +func (k Key) Validate() error { + if strings.IndexByte(k.Table, keySeparator) >= 0 { + return fmt.Errorf("%w: table name %q holds a NUL byte", ErrInvalidKey, k.Table) + } + return nil +} + +// Hashed reports whether k's id is stored as its SHA-256 rather than +// verbatim. +func (k Key) Hashed() bool { + return len(k.ID) > MaxIDBytes || (k.ID != "" && k.ID[0] == hashedID) +} + +// AppendKey appends k's stored form, under the tenant prefix from KeyPrefix, +// to dst. k must be valid. +func AppendKey(dst, prefix []byte, k Key) []byte { + dst = append(dst, prefix...) + dst = append(dst, k.Table...) + dst = append(dst, keySeparator) + if k.Hashed() { + sum := sha256.Sum256([]byte(k.ID)) + dst = append(dst, hashedID) + return append(dst, sum[:]...) + } + return append(dst, k.ID...) +} diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index 3337e610..b078d3e8 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -3,25 +3,38 @@ package dedupe import ( "context" "errors" + "fmt" "sync" + "time" + + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/metric" ) -// ErrDisabled is returned by Managed.CheckAndMark while dedupe is switched -// off. The ingest handler consults the settings snapshot before calling, so +// ErrDisabled is returned by Managed's calls while dedupe is switched off. The ingest handler consults the settings snapshot before calling, so // it only sees this in the window of a reload that flips dedupe.enabled: // the snapshot and the store transition at different instants, and a record // caught between them is published un-deduped rather than failed. var ErrDisabled = errors.New("dedupe is disabled") -// ErrUnavailable is returned by Managed.CheckAndMark when dedupe is switched -// on but the store failed to open. Ingest fails closed on it — the settings -// asked for dedupe, so publishing un-deduped is not a fallback. +// ErrUnavailable is returned by Managed's calls when dedupe is switched on +// but the store failed to open, and wrapped by a backend's error when a +// retry later can succeed. Ingest fails closed on it — the settings asked +// for dedupe, so publishing un-deduped is not a fallback. var ErrUnavailable = errors.New("dedupe store is not open") +// hashedIDCounter counts ids stored as their SHA-256 (Key.Hashed): an id +// longer than MaxIDBytes is a producer sending something other than an id. +var hashedIDCounter, _ = otel.Meter("wavehouse-dedupe").Int64Counter( + "wavehouse_dedupe_hashed_id_total", + metric.WithDescription("Dedupe ids stored as their SHA-256 because they exceed the verbatim length limit"), +) + // Managed is a Deduplicator whose backing store follows the hot-reloadable // dedupe.enabled setting: Apply(true) opens it through the function -// NewManaged was given, Apply(false) closes it, and in-flight CheckAndMark -// calls are serialized against that swap so a reload can never close the +// NewManaged was given, Apply(false) closes it, and in-flight Reserve, Commit +// and Release calls are serialized against that swap so a reload can never close the // store under a lookup. Which store that is — a tenant's share of the // embedded Pebble instance (Embedded.Tenant), a remote backend's view later — // is the opener's business, so every backend gets the same switch semantics. @@ -70,18 +83,101 @@ func (m *Managed) Open() bool { return m.db != nil } -// CheckAndMark delegates to the open store; ErrDisabled while switched off, -// ErrUnavailable while switched on but not open. -func (m *Managed) CheckAndMark(ctx context.Context, eventID string) (bool, error) { +// Reserve checks every key is storable, collapses a key repeated inside keys +// to one backend claim — later occurrences answer Duplicate — and delegates +// the rest to the open store; ErrDisabled while switched off, ErrUnavailable +// while switched on but not open. +func (m *Managed) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { + for _, k := range keys { + if err := k.Validate(); err != nil { + return nil, err + } + if k.Hashed() { + hashedIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", k.Table))) + } + } + m.mu.RLock() + defer m.mu.RUnlock() + if err := m.usable(); err != nil { + return nil, err + } + first := make(map[Key]int, len(keys)) + unique := make([]Key, 0, len(keys)) + for _, k := range keys { + if _, seen := first[k]; !seen { + first[k] = len(unique) + unique = append(unique, k) + } + } + got, err := m.db.Reserve(ctx, unique, lease) + if err != nil { + return nil, err + } + if len(got) != len(unique) { + _ = m.db.Release(context.WithoutCancel(ctx), got) + return nil, fmt.Errorf("dedupe backend answered %d claims for %d keys", len(got), len(unique)) + } + if len(unique) == len(keys) { + return got, nil + } + claims := make([]Claim, len(keys)) + answered := make([]bool, len(unique)) + for i, k := range keys { + j := first[k] + if answered[j] { + claims[i] = Claim{Key: k, Status: Duplicate} + continue + } + answered[j] = true + claims[i] = got[j] + } + return claims, nil +} + +// Commit delegates the Claimed claims to the open store, with Reserve's +// switch semantics. +func (m *Managed) Commit(ctx context.Context, claims []Claim, retention time.Duration) error { + return m.withClaimed(claims, func(db Deduplicator, claimed []Claim) error { + return db.Commit(ctx, claimed, retention) + }) +} + +// Release delegates the Claimed claims to the open store, with Reserve's +// switch semantics. +func (m *Managed) Release(ctx context.Context, claims []Claim) error { + return m.withClaimed(claims, func(db Deduplicator, claimed []Claim) error { + return db.Release(ctx, claimed) + }) +} + +func (m *Managed) withClaimed(claims []Claim, do func(Deduplicator, []Claim) error) error { + claimed := make([]Claim, 0, len(claims)) + for _, c := range claims { + if c.Status == Claimed { + claimed = append(claimed, c) + } + } + if len(claimed) == 0 { + return nil + } m.mu.RLock() defer m.mu.RUnlock() - if !m.enabled { - return false, ErrDisabled + if err := m.usable(); err != nil { + return err } - if m.db == nil { - return false, ErrUnavailable + return do(m.db, claimed) +} + +// usable is the switch's answer: nil when the store may be called. Callers +// hold mu. +func (m *Managed) usable() error { + switch { + case !m.enabled: + return ErrDisabled + case m.db == nil: + return ErrUnavailable } - return m.db.CheckAndMark(ctx, eventID) + return nil } // Close releases the store if open. Safe to call when already closed. diff --git a/internal/dedupe/managed_test.go b/internal/dedupe/managed_test.go index 74904be6..13880157 100644 --- a/internal/dedupe/managed_test.go +++ b/internal/dedupe/managed_test.go @@ -4,6 +4,7 @@ import ( "context" "errors" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -16,28 +17,28 @@ func TestManaged_FollowsEnabled(t *testing.T) { ctx := context.Background() assert.False(t, m.Open()) - _, err := m.CheckAndMark(ctx, "e1") + _, err := mark(ctx, m, "e1") require.ErrorIs(t, err, ErrDisabled) require.NoError(t, m.Apply(true)) require.NoError(t, m.Apply(true), "re-applying the same state is a no-op") assert.True(t, m.Open()) - dup, err := m.CheckAndMark(ctx, "e1") + dup, err := mark(ctx, m, "e1") require.NoError(t, err) assert.False(t, dup) - dup, err = m.CheckAndMark(ctx, "e1") + dup, err = mark(ctx, m, "e1") require.NoError(t, err) assert.True(t, dup) require.NoError(t, m.Apply(false)) require.NoError(t, m.Apply(false)) assert.False(t, m.Open()) - _, err = m.CheckAndMark(ctx, "e1") + _, err = mark(ctx, m, "e1") require.ErrorIs(t, err, ErrDisabled) // Re-enabling reopens the same instance: previously seen ids persist. require.NoError(t, m.Apply(true)) - dup, err = m.CheckAndMark(ctx, "e1") + dup, err = mark(ctx, m, "e1") require.NoError(t, err) assert.True(t, dup, "toggling off and on must not forget seen ids") } @@ -45,34 +46,53 @@ func TestManaged_FollowsEnabled(t *testing.T) { // memDedup is the smallest possible backend: what a shared remote store's // per-tenant view would be, minus the network. type memDedup struct { - seen map[string]bool - closed bool + seen map[Key]bool + closed bool + reserved [][]Key // every Reserve's keys, as the backend saw them + short bool // answer one claim too few } -func (m *memDedup) CheckAndMark(_ context.Context, id string) (bool, error) { - if m.seen[id] { - return true, nil +func (m *memDedup) Reserve(_ context.Context, keys []Key, _ time.Duration) ([]Claim, error) { + m.reserved = append(m.reserved, keys) + claims := make([]Claim, 0, len(keys)) + for _, k := range keys { + st := Claimed + if m.seen[k] { + st = Duplicate + } + claims = append(claims, Claim{Key: k, Status: st, Token: "t"}) } - m.seen[id] = true - return false, nil + if m.short { + claims = claims[1:] + } + return claims, nil } -func (m *memDedup) Close() error { m.closed = true; return nil } + +func (m *memDedup) Commit(_ context.Context, claims []Claim, _ time.Duration) error { + for _, c := range claims { + m.seen[c.Key] = true + } + return nil +} + +func (m *memDedup) Release(context.Context, []Claim) error { return nil } +func (m *memDedup) Close() error { m.closed = true; return nil } // The switch semantics belong to Managed, not to Pebble: any Deduplicator // an opener returns gets them, and a failing opener reads as unavailable. func TestManaged_AnyBackend(t *testing.T) { t.Parallel() ctx := context.Background() - backend := &memDedup{seen: map[string]bool{}} + backend := &memDedup{seen: map[Key]bool{}} m := NewManaged(func() (Deduplicator, error) { return backend, nil }) - _, err := m.CheckAndMark(ctx, "e1") + _, err := mark(ctx, m, "e1") require.ErrorIs(t, err, ErrDisabled) require.NoError(t, m.Apply(true)) - dup, err := m.CheckAndMark(ctx, "e1") + dup, err := mark(ctx, m, "e1") require.NoError(t, err) assert.False(t, dup) - dup, err = m.CheckAndMark(ctx, "e1") + dup, err = mark(ctx, m, "e1") require.NoError(t, err) assert.True(t, dup) require.NoError(t, m.Close()) @@ -81,7 +101,7 @@ func TestManaged_AnyBackend(t *testing.T) { failing := NewManaged(func() (Deduplicator, error) { return nil, errors.New("backend down") }) require.ErrorContains(t, failing.Apply(true), "backend down") assert.False(t, failing.Open()) - _, err = failing.CheckAndMark(ctx, "e1") + _, err = mark(ctx, failing, "e1") require.ErrorIs(t, err, ErrUnavailable) } @@ -90,9 +110,46 @@ func TestManaged_OpenFailureStaysClosed(t *testing.T) { m := NewManaged(func() (Deduplicator, error) { return nil, errors.New("disk full") }) require.Error(t, m.Apply(true)) assert.False(t, m.Open()) - _, err := m.CheckAndMark(context.Background(), "e1") + _, err := mark(context.Background(), m, "e1") require.ErrorIs(t, err, ErrUnavailable, "switched on but not open must fail closed, not read as disabled") require.NoError(t, m.Close()) - _, err = m.CheckAndMark(context.Background(), "e1") + _, err = mark(context.Background(), m, "e1") require.ErrorIs(t, err, ErrDisabled) } + +// Managed collapses a key repeated in one call before the backend sees it, +// so every backend answers repeats alike, and hands the backend only the +// claims it made. +func TestManaged_CollapsesRepeats(t *testing.T) { + t.Parallel() + ctx := context.Background() + backend := &memDedup{seen: map[Key]bool{}} + m := NewManaged(func() (Deduplicator, error) { return backend, nil }) + require.NoError(t, m.Apply(true)) + a, b := Key{Table: "t", ID: "a"}, Key{Table: "t", ID: "b"} + + claims, err := m.Reserve(ctx, []Key{a, b, a}, time.Second) + require.NoError(t, err) + assert.Equal(t, [][]Key{{a, b}}, backend.reserved, "the backend sees each key once") + assert.Equal(t, []Claim{{Key: a, Status: Claimed, Token: "t"}, {Key: b, Status: Claimed, Token: "t"}, {Key: a, Status: Duplicate}}, claims) + + backend.short = true + _, err = m.Reserve(ctx, []Key{a}, time.Second) + require.ErrorContains(t, err, "answered 0 claims for 1 keys", "a backend answering the wrong count is refused, not indexed past") +} + +// Commit and Release follow the switch like Reserve, and a call with no +// Claimed claim never reaches the backend. +func TestManaged_CommitAndReleaseFollowTheSwitch(t *testing.T) { + t.Parallel() + ctx := context.Background() + claimed := []Claim{{Key: Key{Table: "t", ID: "a"}, Status: Claimed, Token: "t"}} + m := NewManaged(func() (Deduplicator, error) { return nil, errors.New("down") }) + require.ErrorIs(t, m.Commit(ctx, claimed, 0), ErrDisabled) + require.ErrorIs(t, m.Release(ctx, claimed), ErrDisabled) + require.NoError(t, m.Commit(ctx, []Claim{{Status: Duplicate}}, 0), "nothing to commit") + + require.Error(t, m.Apply(true)) + require.ErrorIs(t, m.Commit(ctx, claimed, 0), ErrUnavailable) + require.ErrorIs(t, m.Release(ctx, claimed), ErrUnavailable) +} diff --git a/internal/dedupe/stores_test.go b/internal/dedupe/stores_test.go index 15136343..ea2ae065 100644 --- a/internal/dedupe/stores_test.go +++ b/internal/dedupe/stores_test.go @@ -30,7 +30,7 @@ func TestStores_ForBuildsOneClosedStorePerTenant(t *testing.T) { assert.Same(t, acme, s.For("acme"), "one store per tenant, however often it is named") assert.NotSame(t, acme, s.For("globex")) assert.False(t, acme.Open(), "built closed: nothing opens until the tenant's switch is applied") - _, err := acme.CheckAndMark(ctx, "e1") + _, err := mark(ctx, acme, "e1") require.ErrorIs(t, err, ErrDisabled, "a store not yet applied answers as a disabled one, the reload-window case") assert.NoDirExists(t, e.Dir()) @@ -49,12 +49,12 @@ func TestStores_TenantsDoNotShareSeenIDs(t *testing.T) { } for _, id := range tenants { - dup, err := s.For(id).CheckAndMark(ctx, "e1") + dup, err := mark(ctx, s.For(id), "e1") require.NoError(t, err) assert.False(t, dup, "%s: the same event id is first seen in each tenant", id) } for _, id := range tenants { - dup, err := s.For(id).CheckAndMark(ctx, "e1") + dup, err := mark(ctx, s.For(id), "e1") require.NoError(t, err) assert.True(t, dup, "%s: and a duplicate within its own tenant", id) } @@ -67,7 +67,7 @@ func TestStores_RetainClosesTheRestAndKeepsTheirData(t *testing.T) { acme, globex := s.For("acme"), s.For("globex") require.NoError(t, acme.Apply(true)) require.NoError(t, globex.Apply(true)) - _, err := acme.CheckAndMark(ctx, "e1") + _, err := mark(ctx, acme, "e1") require.NoError(t, err) require.NoError(t, s.Retain(func(id tenant.ID) bool { return id == "globex" })) @@ -79,7 +79,7 @@ func TestStores_RetainClosesTheRestAndKeepsTheirData(t *testing.T) { restored := s.For("acme") assert.NotSame(t, acme, restored, "the closed store was forgotten") require.NoError(t, restored.Apply(true)) - dup, err := restored.CheckAndMark(ctx, "e1") + dup, err := mark(ctx, restored, "e1") require.NoError(t, err) assert.True(t, dup, "an id seen before the tenant was dropped is still seen") } @@ -88,8 +88,12 @@ func TestStores_RetainClosesTheRestAndKeepsTheirData(t *testing.T) { // slow I/O, as the last Pebble close waiting on a compaction. type gatedDedup struct{ entered, release chan struct{} } -func (g *gatedDedup) CheckAndMark(context.Context, string) (bool, error) { return false, nil } -func (g *gatedDedup) Close() error { g.entered <- struct{}{}; <-g.release; return nil } +func (g *gatedDedup) Reserve(context.Context, []Key, time.Duration) ([]Claim, error) { + return nil, nil +} +func (g *gatedDedup) Commit(context.Context, []Claim, time.Duration) error { return nil } +func (g *gatedDedup) Release(context.Context, []Claim) error { return nil } +func (g *gatedDedup) Close() error { g.entered <- struct{}{}; <-g.release; return nil } // One tenant's I/O is that tenant's wait alone: Retain edits the map under // the lock and closes outside it, so a dropped tenant's slow close never diff --git a/internal/settings/validate.go b/internal/settings/validate.go index 598c4aa7..b3eb4231 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -451,14 +451,17 @@ func (v *validator) checkIDField(path string, val *string) { } // checkTableName rejects a per-table override key that could never match a -// table: empty, or carrying surrounding whitespace. Shared by the dedupe and -// dlq override maps. +// table: empty, carrying surrounding whitespace, or holding NUL. Shared by +// the dedupe and dlq override maps. func (v *validator) checkTableName(mapPath, table string) { switch { case table == "": v.errorf(FileConfig, mapPath, "table name must not be empty") case strings.TrimSpace(table) != table: v.errorf(FileConfig, mapPath+"."+table, "table name %q has surrounding whitespace", table) + case strings.ContainsRune(table, 0): + // The dedupe key ends the table with NUL (dedupe.KeyPrefix). + v.errorf(FileConfig, mapPath+"."+table, "table name %q holds a NUL byte", table) } } diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index 56be6b2a..4dae9c9d 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -264,6 +264,7 @@ func TestValidate_ContentRules(t *testing.T) { {"padded override id_field", FileConfig, `{"dedupe": {"tables": {"clicks": {"id_field": "click_id "}}}}`, "dedupe.tables.clicks.id_field"}, {"empty override table name", FileConfig, `{"dedupe": {"tables": {"": {"id_field": "x"}}}}`, "table name must not be empty"}, {"override table whitespace", FileConfig, `{"dedupe": {"tables": {" clicks": {"require_id": true}}}}`, "surrounding whitespace"}, + {"override table NUL", FileConfig, `{"dedupe": {"tables": {"cli\u0000cks": {"require_id": true}}}}`, "holds a NUL byte"}, {"empty override id_field", FileConfig, `{"dedupe": {"tables": {"clicks": {"id_field": ""}}}}`, "dedupe.tables.clicks.id_field: must not be empty"}, {"negative max rows", FileConfig, `{"query": {"default_max_rows": -1}}`, "must be >= 1"}, {"zero max rows", FileConfig, `{"query": {"default_max_rows": 0}}`, "must be >= 1"}, diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index 44f3ebe6..cdb1e33e 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -3,6 +3,7 @@ package testutil import ( "bytes" "context" + "fmt" "io" "net/http" "sync" @@ -10,6 +11,7 @@ import ( "time" "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/dedupe" "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -33,6 +35,10 @@ type MockPublisher struct { mu sync.Mutex Messages []PublishedMessage Err error // if set, Publish and DeadLetter return this error + // ErrAfter lets that many calls succeed before Err applies, to fail a + // batch part-way through. + ErrAfter int + calls int } // PublishedMessage records a single Publish or DeadLetter call, with the @@ -53,15 +59,16 @@ func (m *MockPublisher) DeadLetter(_ context.Context, msg *mq.Message, opts ...m } func (m *MockPublisher) record(pm PublishedMessage, opts []mq.PublishOpt) error { - if m.Err != nil { + m.mu.Lock() + defer m.mu.Unlock() + m.calls++ + if m.Err != nil && m.calls > m.ErrAfter { return m.Err } headers := mq.Headers{} for _, opt := range opts { opt(headers) } - m.mu.Lock() - defer m.mu.Unlock() pm.Headers = headers m.Messages = append(m.Messages, pm) return nil @@ -104,28 +111,92 @@ func (m *MockSubscriber) Close() error { return nil } // ── Mock Deduplicator ──────────────────────────────────────────── -// MockDeduplicator implements dedupe.Deduplicator for testing. +// MockDeduplicator implements dedupe.Deduplicator in memory, with per-phase +// error injection and a record of what was committed and released. type MockDeduplicator struct { - mu sync.Mutex - seen map[string]bool - Err error // if set, CheckAndMark returns this error + mu sync.Mutex + committed map[dedupe.Key]bool + pending map[dedupe.Key]string + tokens int + // Err, if set, fails Reserve; CommitErr and ReleaseErr fail their phase. + Err error + CommitErr error + ReleaseErr error + Released []dedupe.Claim // every claim Release was given } +var _ dedupe.Deduplicator = (*MockDeduplicator)(nil) + func NewMockDeduplicator() *MockDeduplicator { - return &MockDeduplicator{seen: make(map[string]bool)} + return &MockDeduplicator{committed: map[dedupe.Key]bool{}, pending: map[dedupe.Key]string{}} } -func (m *MockDeduplicator) CheckAndMark(_ context.Context, eventID string) (bool, error) { +func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time.Duration) ([]dedupe.Claim, error) { if m.Err != nil { - return false, m.Err + return nil, m.Err + } + m.mu.Lock() + defer m.mu.Unlock() + claims := make([]dedupe.Claim, 0, len(keys)) + for _, k := range keys { + switch { + case m.committed[k]: + claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.Duplicate}) + case m.pending[k] != "": + claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.InFlight}) + default: + m.tokens++ + tok := fmt.Sprint(m.tokens) + m.pending[k] = tok + claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.Claimed, Token: tok}) + } + } + return claims, nil +} + +func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ time.Duration) error { + if m.CommitErr != nil { + return m.CommitErr + } + m.mu.Lock() + defer m.mu.Unlock() + for _, c := range claims { + if c.Status == dedupe.Claimed { + m.committed[c.Key] = true + delete(m.pending, c.Key) + } } + return nil +} + +func (m *MockDeduplicator) Release(_ context.Context, claims []dedupe.Claim) error { m.mu.Lock() defer m.mu.Unlock() - if m.seen[eventID] { - return true, nil + m.Released = append(m.Released, claims...) + if m.ReleaseErr != nil { + return m.ReleaseErr } - m.seen[eventID] = true - return false, nil + for _, c := range claims { + if c.Status == dedupe.Claimed && m.pending[c.Key] == c.Token { + delete(m.pending, c.Key) + } + } + return nil +} + +// Hold claims k as another in-flight request would, so Reserve answers +// InFlight for it. +func (m *MockDeduplicator) Hold(k dedupe.Key) { + m.mu.Lock() + defer m.mu.Unlock() + m.pending[k] = "held" +} + +// Pending reports whether k is claimed and neither committed nor released. +func (m *MockDeduplicator) Pending(k dedupe.Key) bool { + m.mu.Lock() + defer m.mu.Unlock() + return m.pending[k] != "" } func (m *MockDeduplicator) Close() error { return nil } From 0ec037e5e15b6f3addaba50eeb363d1436ee7e92 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:50:18 -0400 Subject: [PATCH 017/108] fix(dedupe): state what Managed guarantees a backend; default a zero lease Also document the upgrade, the SDK's new 503 cause, and the release on a failed publish. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 ++-- docs/src/content/docs/architecture.md | 6 +++--- docs/src/content/docs/deployment.md | 6 +++++- docs/src/content/docs/sdk/reference.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/dedupe/dedupe.go | 9 ++++++--- internal/dedupe/managed.go | 7 +++++-- internal/dedupe/managed_test.go | 10 ++++++++-- 9 files changed, 32 insertions(+), 16 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 038096e1..869dfe42 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,7 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble` until a later sweep drops them. A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble` until a later sweep drops them ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 9040a6bb..cc89c314 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -273,7 +273,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error | +| 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate. | | 503 | `{"error":"service unavailable"}` | NATS JetStream stream full (backpressure). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | @@ -385,7 +385,7 @@ A `200` is returned whenever the body was read and the records were processed | 403 | `{"error":"forbidden"}` (empty-role variant: `forbidden: request has no role and no public default_role is configured`) | The resolved role lacks `insert` on the table (checked once, before any record) | | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | -| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch | +| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest | | 503 | `{"error":"service unavailable"}` | NATS JetStream full (backpressure) mid-batch; includes `Retry-After: 30` | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 19d25bed..3e4a9a86 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -124,8 +124,8 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. -- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes a window of claims in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key and collapses a key repeated in one call before the backend sees it, once for every backend. +- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 63cff252..799df599 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -389,7 +389,7 @@ The folder name is the tenant id, and each folder is a complete settings directo **The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. +**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant and table), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it, and in each table. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. **What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. @@ -417,6 +417,10 @@ ORDER BY (page); WaveHouse discovers this schema on startup and refreshes it every `schema.refresh_interval` seconds (settings directory; seed default 60). You can also trigger an immediate refresh via `POST /v1/ops/schema/refresh` (admin-only). +## Upgrading across the dedupe key change + +The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread, until a later sweep removes them. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. + ## Upgrading across the v2 ingest envelope The NATS envelope changed shape in this release: the row now travels positionally, with `format`, `columns` and `row` replacing `data`. **The new worker cannot read a message published by an older version** — it carries no `format`, so there is no way to say which value belongs to which column. diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index af0cddef..a7a44387 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -32,7 +32,7 @@ The SDK **never throws** for anything the server returns — all API errors come | 403 | `HTTP_403` | No | Insufficient permissions | | 404 | `HTTP_404` | No | Table, pipe, or tenant not found | | 500 | `HTTP_500` | Yes | Server error (retried per `maxRetries`) | -| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, or a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on that last cause waits the 30 s; a stream re-dials on its own jittered backoff instead | +| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | | 0 | `NETWORK_ERROR` | Yes | Network failure (retried with exponential backoff) | | 0 | `ABORTED` | No | Request canceled via `AbortSignal` | | 0 | `SSE_CONNECT_ERROR` | No | Stream could not be started (e.g. a non-absolute `baseURL`) | diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index e8c31e15..b2c39856 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/dedupe/dedupe.go b/internal/dedupe/dedupe.go index 7602443a..08020265 100644 --- a/internal/dedupe/dedupe.go +++ b/internal/dedupe/dedupe.go @@ -27,8 +27,8 @@ const ( // after the lease. Claimed Status = iota + 1 // Duplicate means the key was committed earlier and has not expired: skip - // the record. Also returned for a key repeated inside one Reserve call, - // after its first occurrence. + // the record. Managed also answers it for a key repeated inside one + // Reserve call, after its first occurrence, whatever the first answered. Duplicate // InFlight means another request holds a live claim on the key. Its // outcome is not known yet, so the caller answers 503 and the client @@ -57,7 +57,10 @@ type Claim struct { Token string } -// Deduplicator is a tenant's store of seen ids. +// Deduplicator is a tenant's store of seen ids. Callers reach every backend +// through Managed, which hands a backend distinct, valid keys, a lease > 0, +// and only Claimed claims to Commit and Release — a backend may assume all +// three, and Managed's callers get the behaviour below either way. // // Reserve is atomic per key: of any number of concurrent Reserves for the // same key — in this process or any other sharing the backend — at most one diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index b078d3e8..1fc3a7d5 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -84,8 +84,8 @@ func (m *Managed) Open() bool { } // Reserve checks every key is storable, collapses a key repeated inside keys -// to one backend claim — later occurrences answer Duplicate — and delegates -// the rest to the open store; ErrDisabled while switched off, ErrUnavailable +// to one backend claim — later occurrences answer Duplicate — reads a lease +// <= 0 as DefaultLease, and delegates the rest to the open store; ErrDisabled while switched off, ErrUnavailable // while switched on but not open. func (m *Managed) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { for _, k := range keys { @@ -96,6 +96,9 @@ func (m *Managed) Reserve(ctx context.Context, keys []Key, lease time.Duration) hashedIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", k.Table))) } } + if lease <= 0 { + lease = DefaultLease + } m.mu.RLock() defer m.mu.RUnlock() if err := m.usable(); err != nil { diff --git a/internal/dedupe/managed_test.go b/internal/dedupe/managed_test.go index 13880157..97432f8e 100644 --- a/internal/dedupe/managed_test.go +++ b/internal/dedupe/managed_test.go @@ -49,11 +49,13 @@ type memDedup struct { seen map[Key]bool closed bool reserved [][]Key // every Reserve's keys, as the backend saw them - short bool // answer one claim too few + leases []time.Duration + short bool // answer one claim too few } -func (m *memDedup) Reserve(_ context.Context, keys []Key, _ time.Duration) ([]Claim, error) { +func (m *memDedup) Reserve(_ context.Context, keys []Key, lease time.Duration) ([]Claim, error) { m.reserved = append(m.reserved, keys) + m.leases = append(m.leases, lease) claims := make([]Claim, 0, len(keys)) for _, k := range keys { st := Claimed @@ -133,6 +135,10 @@ func TestManaged_CollapsesRepeats(t *testing.T) { assert.Equal(t, [][]Key{{a, b}}, backend.reserved, "the backend sees each key once") assert.Equal(t, []Claim{{Key: a, Status: Claimed, Token: "t"}, {Key: b, Status: Claimed, Token: "t"}, {Key: a, Status: Duplicate}}, claims) + _, err = m.Reserve(ctx, []Key{{Table: "t", ID: "c"}}, 0) + require.NoError(t, err) + assert.Equal(t, []time.Duration{time.Second, DefaultLease}, backend.leases, "no lease is the default, never an already-lapsed claim") + backend.short = true _, err = m.Reserve(ctx, []Key{a}, time.Second) require.ErrorContains(t, err, "answered 0 claims for 1 keys", "a backend answering the wrong count is refused, not indexed past") From 9344f1b8cd33b50dc3534fdb0607b0a201df973a Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 23:53:39 -0400 Subject: [PATCH 018/108] fix(mq): pace a park's reopen, and warn only on a missing consumer; review fixes --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 +-- internal/ingest/sweeper.go | 18 +++++++++- internal/ingest/sweeper_test.go | 27 ++++++++++++++ internal/mq/embedded.go | 51 ++++++++++++++------------- internal/mq/embedded_test.go | 23 ++++++------ internal/mq/mq.go | 7 ++-- 7 files changed, 90 insertions(+), 42 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 27cbd578..70ce51e1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fb6f03fc..fe0c94c9 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -145,10 +145,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. -- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. +- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/internal/ingest/sweeper.go b/internal/ingest/sweeper.go index b0d4a2d1..4024b9de 100644 --- a/internal/ingest/sweeper.go +++ b/internal/ingest/sweeper.go @@ -61,7 +61,7 @@ func (s *Sweeper) sweep(ctx context.Context) { } _, err := s.purger.PurgeAcked(ctx, BufferConsumerName, cutoffs) if err != nil { - if errors.Is(err, mq.ErrConsumerNotFound) { + if onlyConsumerNotFound(err) { // Consumer may not exist yet if no messages have been ingested. slog.WarnContext(ctx, "sweeper: buffer consumer not found (may not exist yet)", "error", err) return @@ -69,3 +69,19 @@ func (s *Sweeper) sweep(ctx context.Context) { slog.ErrorContext(ctx, "sweeper: purge", "error", err) } } + +// onlyConsumerNotFound reports whether every tenant's failure err joins is a +// missing buffer consumer — the one failure expected before the worker has +// created it. Any other failure among them keeps the sweep's report at +// ERROR: a tenant whose purge keeps failing fills toward its budget. +func onlyConsumerNotFound(err error) bool { + if joined, ok := err.(interface{ Unwrap() []error }); ok { + for _, e := range joined.Unwrap() { + if !onlyConsumerNotFound(e) { + return false + } + } + return true + } + return errors.Is(err, mq.ErrConsumerNotFound) +} diff --git a/internal/ingest/sweeper_test.go b/internal/ingest/sweeper_test.go index 9b1eabab..3d585561 100644 --- a/internal/ingest/sweeper_test.go +++ b/internal/ingest/sweeper_test.go @@ -3,12 +3,15 @@ package ingest import ( "context" "errors" + "fmt" + "log/slog" "testing" "time" "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil" + "github.com/Wave-RF/WaveHouse/internal/testutil/logtest" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" ) @@ -63,6 +66,30 @@ func TestSweep_ErrorsDoNotPanic(t *testing.T) { } } +// A missing buffer consumer is the expected failure, before the worker has +// created it, and only a warning; any other tenant's failure in the same +// sweep — the purger joins one per tenant — keeps the report at ERROR. +func TestSweep_OnlyAMissingConsumerIsAWarning(t *testing.T) { + missing := fmt.Errorf("tenant acme: %w", mq.ErrConsumerNotFound) + for _, tt := range []struct { + name string + err error + want, not string + }{ + {"a missing consumer", errors.Join(missing), "WARN", "ERROR"}, + {"a missing consumer beside another failure", errors.Join(missing, errors.New("tenant globex: get stream: stream not found")), "ERROR", "WARN"}, + {"another failure", errors.New("broker unavailable"), "ERROR", "WARN"}, + } { + t.Run(tt.name, func(t *testing.T) { + logs := logtest.Capture(t, slog.LevelDebug) + s := NewSweeper(&testutil.MockPurger{Err: tt.err}, func() map[tenant.ID]time.Duration { return nil }) + s.sweep(context.Background()) + assert.Contains(t, logs.String(), `"level":"`+tt.want+`"`) + assert.NotContains(t, logs.String(), `"level":"`+tt.not+`"`) + }) + } +} + // --------------------------------------------------------------------------- // Start() context cancellation test // --------------------------------------------------------------------------- diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 32d5047d..830be900 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -73,17 +73,17 @@ type EmbeddedNATS struct { // leave behind a stream JetStream goes on to create, which no consumer // holds. Written under mu, read without it. opened sync.Map // tenant.ID → struct{} - // reopening merges into one attempt the publishes that find the same - // tenant's queue not open, and failedOpen holds, for a tenant whose last - // such attempt failed, its error and until when its publishes take that - // as their answer (openForPublish). + // reopening merges into one attempt the publishes and parks that find + // the same tenant's queue not open, and failedOpen holds, for a tenant + // whose last such attempt failed, its error and until when its publishes + // and parks take that as their answer (reopenPaced). reopening singleflight.Group failedOpen sync.Map // tenant.ID → openFailure } -// openFailure is a publish's failed attempt to open a tenant's queue, and -// until when the tenant's publishes are refused with its error rather than -// trying again. +// openFailure is a publish's or park's failed attempt to open a tenant's +// queue, and until when the tenant's publishes and parks are refused with its +// error rather than trying again. type openFailure struct { until time.Time err error @@ -126,9 +126,9 @@ const ( // resizeTimeouts when it opens a queue: the consumers join on a budget of // their own (apply). rollbackTimeout = 5 * time.Second - // publishRetry is how long a tenant's publishes are refused at once after - // one failed to open its queue (openForPublish). - publishRetry = 5 * time.Second + // reopenRetry is how long a tenant's publishes and parks are refused at + // once after one failed to open its queue (reopenPaced). + reopenRetry = 5 * time.Second ) // errNoQueue is why a publish or park finds no queue it can open: no budget @@ -511,7 +511,7 @@ func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { // subject). A tenant with no queue has one opened at the budget last asked // for it (see SetMaxBytes) — and so does one whose stream exists but whose // queue the broker has not recorded open, since no consumer may hold that -// stream (see openForPublish for how often a publish tries). A queue that +// stream (see reopenPaced for how often a publish tries). A queue that // cannot be opened — none asked for yet, or JetStream refused it — and a // queue at its byte budget (DiscardNew) are reported as ErrQueueFull: either // way the tenant's queue takes nothing now, and a retry is the caller's @@ -522,13 +522,13 @@ func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, op return err } if _, ok := e.opened.Load(topic.Tenant); !ok { - if openErr := e.openForPublish(ctx, topic.Tenant); openErr != nil { + if openErr := e.reopenPaced(ctx, topic.Tenant); openErr != nil { return fmt.Errorf("%w: %w", ErrQueueFull, openErr) } } err = e.publish(ctx, subj, data, opts) if errors.Is(err, jetstream.ErrNoStreamResponse) { - if openErr := e.openForPublish(ctx, topic.Tenant); openErr != nil { + if openErr := e.reopenPaced(ctx, topic.Tenant); openErr != nil { return fmt.Errorf("%w: %w", ErrQueueFull, openErr) } err = e.publish(ctx, subj, data, opts) @@ -541,15 +541,15 @@ func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, op return err } -// openForPublish opens tenant id's queue for a publish that found it not open -// (reopen). The publishes that find it so at the same time share one -// attempt, and after an attempt fails the tenant's publishes get its error at -// once, without taking mu, until publishRetry has passed: under clients -// retrying, a queue that cannot open would otherwise hold mu for attempt -// after attempt, and every other tenant's open, resize and reload waits on -// mu. A reload that applies the tenant's budget retries it regardless -// (SetMaxBytes). -func (e *EmbeddedNATS) openForPublish(ctx context.Context, id tenant.ID) error { +// reopenPaced opens tenant id's queue for a publish or park that found it not +// open (reopen). The callers that find it so at the same time share one +// attempt, and after an attempt fails the tenant's publishes and parks get +// its error at once, without taking mu, until reopenRetry has passed: under +// clients retrying, or the worker parking row after row, a queue that cannot +// open would otherwise hold mu for attempt after attempt, and every other +// tenant's open, resize and reload waits on mu. A reload that applies the +// tenant's budget retries it regardless (SetMaxBytes). +func (e *EmbeddedNATS) reopenPaced(ctx context.Context, id tenant.ID) error { if v, ok := e.failedOpen.Load(id); ok { if f := v.(openFailure); time.Now().Before(f.until) { return f.err @@ -558,7 +558,7 @@ func (e *EmbeddedNATS) openForPublish(ctx context.Context, id tenant.ID) error { _, err, _ := e.reopening.Do(string(id), func() (any, error) { err := e.reopen(ctx, id) if err != nil { - e.failedOpen.Store(id, openFailure{until: time.Now().Add(publishRetry), err: err}) + e.failedOpen.Store(id, openFailure{until: time.Now().Add(reopenRetry), err: err}) } return nil, err }) @@ -570,13 +570,14 @@ func (e *EmbeddedNATS) openForPublish(ctx context.Context, id tenant.ID) error { // for the dead-letter one, nothing decoded or re-encoded. The dead-letter // stream is DiscardOld, so a full one drops its oldest parked rows rather than // refusing. A dead-letter stream found missing is opened again with its -// tenant's queue, as Publish does. +// tenant's queue, paced as Publish's is (reopenPaced); a park refused leaves +// its row unacked, to be redelivered. func (e *EmbeddedNATS) DeadLetter(ctx context.Context, msg *Message, opts ...PublishOpt) error { subj := dlqPrefix + msg.topicKey err := e.publish(ctx, subj, msg.Data, opts) if errors.Is(err, jetstream.ErrNoStreamResponse) { if id, ok := keyTenant(msg.topicKey); ok { - if err = e.reopen(ctx, id); err == nil { + if err = e.reopenPaced(ctx, id); err == nil { err = e.publish(ctx, subj, msg.Data, opts) } } diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 85fd5719..6e87ff7b 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -460,12 +460,13 @@ func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) } -// After a publish fails to open its tenant's queue, the tenant's publishes -// are refused at once, without waiting on the broker's lock, until -// publishRetry has passed: under clients retrying, one tenant's broken queue -// would otherwise hold the lock that every other tenant's open, resize and -// reload takes. Once the window has passed, a publish tries again. -func TestEmbeddedNATS_Publish_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T) { +// After a publish fails to open its tenant's queue, the tenant's publishes — +// and its parks, which find the dead-letter stream missing — are refused at +// once, without waiting on the broker's lock, until reopenRetry has passed: +// under clients retrying, or the worker parking row after row, one tenant's +// broken queue would otherwise hold the lock that every other tenant's open, +// resize and reload takes. Once the window has passed, a publish tries again. +func TestEmbeddedNATS_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T) { dir := t.TempDir() block := filepath.Join(dir, "jetstream", "$G", "streams", dlqStreamName("acme")) obstruct := func() { @@ -486,14 +487,15 @@ func TestEmbeddedNATS_Publish_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T require.ErrorIs(t, err, os.ErrNotExist) } - // The queue could open now, but within the window a publish tries - // nothing: it is refused while the lock is held elsewhere. + // The queue could open now, but within the window a publish or park + // tries nothing: each is refused while the lock is held elsewhere. e.mu.Lock() - var paced error + var paced, parked error done := make(chan struct{}) go func() { defer close(done) paced = e.Publish(ctx, acme, []byte("x")) + parked = e.DeadLetter(ctx, NewMessage(ctx, acme, []byte("x"), time.Now(), nil, nil, nil)) }() var returned bool select { @@ -503,8 +505,9 @@ func TestEmbeddedNATS_Publish_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T } e.mu.Unlock() <-done - require.True(t, returned, "a paced publish waited on the broker's lock") + require.True(t, returned, "a paced publish or park waited on the broker's lock") require.ErrorIs(t, paced, ErrQueueFull) + require.Error(t, parked) assert.Zero(t, e.MaxBytes("acme")) v, ok := e.failedOpen.Load(tenant.ID("acme")) diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 6626cfa8..3f1c45c1 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -279,9 +279,10 @@ type Purger interface { // olderThan. Either bound alone keeps the event: unacked events are not // yet written, and recent ones are still needed for replay. A tenant // olderThan does not name keeps no history: everything it has - // acknowledged goes. Reports whether anything was removed. - // ErrConsumerNotFound when the consumer has not been created on some - // tenant's queue; the other tenants' are purged all the same. + // acknowledged goes. Reports whether anything was removed, and joins + // each failed tenant's error — ErrConsumerNotFound for one whose queue the + // consumer has not been created on; the other tenants' are purged all the + // same. PurgeAcked(ctx context.Context, consumer string, olderThan map[tenant.ID]time.Time) (purged bool, err error) } From f484511f4f25afba4f6f9ae82404c7682ddb54d6 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:54:31 -0400 Subject: [PATCH 019/108] fix(dedupe): release only claimed claims on a wrong-count answer Also rewrap Managed's doc comments, point the NUL-table check at AppendKey, and describe the two-phase dedupe in architecture's request flow. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/architecture.md | 10 +++++-- internal/dedupe/managed.go | 40 +++++++++++++++++---------- internal/dedupe/managed_test.go | 15 +++++++--- internal/settings/validate.go | 3 +- 4 files changed, 45 insertions(+), 23 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3e4a9a86..aeb9c155 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -226,11 +226,15 @@ Client POST /v1/ingest?table={table} → Canonicalize top-level DateTime/DateTime64 column values to RFC 3339 UTC (rewrites the payload so every consumer shares one spelling; fail-open — an unparseable value passes through verbatim for ClickHouse's parser to judge) - → Optional deduplication check (configurable ID field; a row missing that - field is published un-deduped + logged/counted, or rejected under require_id) + → Optional dedupe: resolve the id (configurable ID field; a row missing it or + setting it to null is published un-deduped + logged/counted, or rejected + under require_id); once the record is encoded, reserve (tenant, table, id): + a duplicate is skipped, an id another request holds → 503 + Retry-After + (the 30s lease) → Publish to NATS JetStream (ingest.{tenant}.{table}) + → Commit the reserved id; on a failed publish, release it instead → 200 OK returned immediately - → (If NATS stream is full: 503 + Retry-After header) + → (If NATS stream is full: 503 + Retry-After header, the id released) Ingest worker pipeline (StartIngestWorker): ← JetStream pull consumer (buffer-consumer) on ingest.> diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index 1fc3a7d5..041487cb 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -12,9 +12,10 @@ import ( "go.opentelemetry.io/otel/metric" ) -// ErrDisabled is returned by Managed's calls while dedupe is switched off. The ingest handler consults the settings snapshot before calling, so -// it only sees this in the window of a reload that flips dedupe.enabled: -// the snapshot and the store transition at different instants, and a record +// ErrDisabled is returned by Managed's calls while dedupe is switched off. +// The ingest handler consults the settings snapshot before calling, so it +// only sees this in the window of a reload that flips dedupe.enabled: the +// snapshot and the store transition at different instants, and a record // caught between them is published un-deduped rather than failed. var ErrDisabled = errors.New("dedupe is disabled") @@ -33,9 +34,9 @@ var hashedIDCounter, _ = otel.Meter("wavehouse-dedupe").Int64Counter( // Managed is a Deduplicator whose backing store follows the hot-reloadable // dedupe.enabled setting: Apply(true) opens it through the function -// NewManaged was given, Apply(false) closes it, and in-flight Reserve, Commit -// and Release calls are serialized against that swap so a reload can never close the -// store under a lookup. Which store that is — a tenant's share of the +// NewManaged was given, Apply(false) closes it, and in-flight Reserve, +// Commit and Release calls are serialized against that swap so a reload can +// never close the store under a lookup. Which store that is — a tenant's share of the // embedded Pebble instance (Embedded.Tenant), a remote backend's view later — // is the opener's business, so every backend gets the same switch semantics. type Managed struct { @@ -85,8 +86,9 @@ func (m *Managed) Open() bool { // Reserve checks every key is storable, collapses a key repeated inside keys // to one backend claim — later occurrences answer Duplicate — reads a lease -// <= 0 as DefaultLease, and delegates the rest to the open store; ErrDisabled while switched off, ErrUnavailable -// while switched on but not open. +// <= 0 as DefaultLease, and delegates the rest to the open store; +// ErrDisabled while switched off, ErrUnavailable while switched on but not +// open. func (m *Managed) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { for _, k := range keys { if err := k.Validate(); err != nil { @@ -117,7 +119,9 @@ func (m *Managed) Reserve(ctx context.Context, keys []Key, lease time.Duration) return nil, err } if len(got) != len(unique) { - _ = m.db.Release(context.WithoutCancel(ctx), got) + if claimed := claimedOnly(got); len(claimed) > 0 { + _ = m.db.Release(context.WithoutCancel(ctx), claimed) + } return nil, fmt.Errorf("dedupe backend answered %d claims for %d keys", len(got), len(unique)) } if len(unique) == len(keys) { @@ -154,12 +158,7 @@ func (m *Managed) Release(ctx context.Context, claims []Claim) error { } func (m *Managed) withClaimed(claims []Claim, do func(Deduplicator, []Claim) error) error { - claimed := make([]Claim, 0, len(claims)) - for _, c := range claims { - if c.Status == Claimed { - claimed = append(claimed, c) - } - } + claimed := claimedOnly(claims) if len(claimed) == 0 { return nil } @@ -171,6 +170,17 @@ func (m *Managed) withClaimed(claims []Claim, do func(Deduplicator, []Claim) err return do(m.db, claimed) } +// claimedOnly is the claims a backend's Commit and Release may be handed. +func claimedOnly(claims []Claim) []Claim { + out := make([]Claim, 0, len(claims)) + for _, c := range claims { + if c.Status == Claimed { + out = append(out, c) + } + } + return out +} + // usable is the switch's answer: nil when the store may be called. Callers // hold mu. func (m *Managed) usable() error { diff --git a/internal/dedupe/managed_test.go b/internal/dedupe/managed_test.go index 97432f8e..944544e9 100644 --- a/internal/dedupe/managed_test.go +++ b/internal/dedupe/managed_test.go @@ -50,6 +50,7 @@ type memDedup struct { closed bool reserved [][]Key // every Reserve's keys, as the backend saw them leases []time.Duration + released []Claim short bool // answer one claim too few } @@ -77,8 +78,11 @@ func (m *memDedup) Commit(_ context.Context, claims []Claim, _ time.Duration) er return nil } -func (m *memDedup) Release(context.Context, []Claim) error { return nil } -func (m *memDedup) Close() error { m.closed = true; return nil } +func (m *memDedup) Release(_ context.Context, claims []Claim) error { + m.released = append(m.released, claims...) + return nil +} +func (m *memDedup) Close() error { m.closed = true; return nil } // The switch semantics belong to Managed, not to Pebble: any Deduplicator // an opener returns gets them, and a failing opener reads as unavailable. @@ -140,8 +144,11 @@ func TestManaged_CollapsesRepeats(t *testing.T) { assert.Equal(t, []time.Duration{time.Second, DefaultLease}, backend.leases, "no lease is the default, never an already-lapsed claim") backend.short = true - _, err = m.Reserve(ctx, []Key{a}, time.Second) - require.ErrorContains(t, err, "answered 0 claims for 1 keys", "a backend answering the wrong count is refused, not indexed past") + backend.seen[b] = true + _, err = m.Reserve(ctx, []Key{{Table: "t", ID: "d"}, b, {Table: "t", ID: "e"}}, time.Second) + require.ErrorContains(t, err, "answered 2 claims for 3 keys", "a backend answering the wrong count is refused, not indexed past") + assert.Equal(t, []Claim{{Key: Key{Table: "t", ID: "e"}, Status: Claimed, Token: "t"}}, backend.released, + "and gets back only the claims it made, never its Duplicate") } // Commit and Release follow the switch like Reserve, and a call with no diff --git a/internal/settings/validate.go b/internal/settings/validate.go index b3eb4231..08575c95 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -460,7 +460,8 @@ func (v *validator) checkTableName(mapPath, table string) { case strings.TrimSpace(table) != table: v.errorf(FileConfig, mapPath+"."+table, "table name %q has surrounding whitespace", table) case strings.ContainsRune(table, 0): - // The dedupe key ends the table with NUL (dedupe.KeyPrefix). + // A NUL separates the dedupe key's fields (dedupe.AppendKey), so a + // table holding one could never be deduped. v.errorf(FileConfig, mapPath+"."+table, "table name %q holds a NUL byte", table) } } From 386ce74a32d907417787f332072c8cf4c2b15d2f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:04:21 -0400 Subject: [PATCH 020/108] docs(dedupe): link the old-key sweep to #220 rather than promise it Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/deployment.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index d6fa914d..56aa5a3a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,7 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble` until a later sweep drops them ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index b297cca0..bd47e10a 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -419,7 +419,7 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres ## Upgrading across the dedupe key change -The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread, until a later sweep removes them. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread; nothing removes them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep that will). Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. ## Upgrading across the v2 ingest envelope From 108499f4158a2820996f39b8e5688c6783c4a4db Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:04:55 -0400 Subject: [PATCH 021/108] feat(dedupe): DynamoDB backend, conformance-tested on dynamodb-local One shared table for every tenant, pk (binary) = the dedupe key, no sort key. Reserve is a conditional PutItem per key (ALL_OLD on failure answers Duplicate or InFlight without a read), Commit a BatchWriteItem with unprocessed-item retries, Release a DeleteItem conditional on the token. An item whose ex has passed is absent to Reserve whether or not TTL has deleted it. Throttles, server faults, timeouts and connection errors wrap ErrUnavailable; a breaker short-circuits Reserve after five in a second. Constructible and tested, not yet selectable at boot (F5). Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 1 + docs/src/content/docs/architecture.md | 3 +- docs/src/content/docs/deployment.md | 72 +++ go.mod | 16 + go.sum | 32 ++ internal/dedupe/dynamodb.go | 605 ++++++++++++++++++++++ internal/dedupe/dynamodb_bench_test.go | 77 +++ internal/dedupe/dynamodb_test.go | 357 +++++++++++++ tests/integration/dedupe_dynamodb_test.go | 302 +++++++++++ tests/integration/setup_test.go | 34 ++ 11 files changed, 1499 insertions(+), 2 deletions(-) create mode 100644 internal/dedupe/dynamodb.go create mode 100644 internal/dedupe/dynamodb_bench_test.go create mode 100644 internal/dedupe/dynamodb_test.go create mode 100644 tests/integration/dedupe_dynamodb_test.go diff --git a/AGENTS.md b/AGENTS.md index 63615686..84b0a5ee 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch); `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims, built and conformance-tested against dynamodb-local but not yet selectable at boot), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index 56aa5a3a..f31bb3c1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 74066932..3cee2210 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -58,7 +58,7 @@ internal/ ├── chconn/ One ClickHouse pool per connection tuple among the served tenants, reconciled on reload under the ceiling ├── chsql/ Shared ClickHouse SQL helpers (identifier quoting, bind-safety) ├── config/ YAML + env var configuration loading -├── dedupe/ Optional deduplication (Reserve/Commit/Release; Pebble) +├── dedupe/ Optional deduplication (Reserve/Commit/Release; Pebble, DynamoDB) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) @@ -125,6 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims within a second short-circuit `Reserve` for a second. `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index bd47e10a..2f97ad13 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -421,6 +421,78 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread; nothing removes them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep that will). Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +## A shared dedupe table on DynamoDB + +:::note[Not selectable yet] +The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. The `dedupe.backend` boot key lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. +::: + +Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend's `create_table` switch is refused unless an `endpoint` override is set, so it only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). + +What the backend requires of the table: + +| Attribute | Type | Role | +|---|---|---| +| `pk` | Binary | Partition key, and the only key: tenant, table and id. No sort key. | +| `st` | Number | `1` = pending claim, `2` = committed. | +| `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | +| `tk` | Binary | The claim token that `Release` matches. | + +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet. Without TTL, though, expired items are never removed and storage keeps growing. The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. + +An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: + +```hcl +resource "aws_dynamodb_table" "wavehouse_dedupe" { + name = "wavehouse-dedupe-${var.environment}" + billing_mode = "PAY_PER_REQUEST" # provisioned + auto scaling once traffic is steady + hash_key = "pk" + deletion_protection_enabled = true + + attribute { + name = "pk" + type = "B" + } + + ttl { + attribute_name = "ex" + enabled = true + } + + server_side_encryption { + enabled = true + } + + tags = { + Name = "wavehouse-dedupe-${var.environment}" + Project = "wavehouse-cloud" + Environment = var.environment # prod | dev | ci | demo | benchmark + ManagedBy = "wavehouse-cloud/infra/stacks/prod-platform" + CostCenter = "data-plane" + } +} + +# The pods' role (EKS Pod Identity or IRSA). No Scan, no CreateTable. +data "aws_iam_policy_document" "wavehouse_dedupe" { + statement { + actions = [ + "dynamodb:PutItem", + "dynamodb:DeleteItem", + "dynamodb:BatchWriteItem", + "dynamodb:DescribeTable", + "dynamodb:DescribeTimeToLive", + ] + resources = [aws_dynamodb_table.wavehouse_dedupe.arn] + } +} +``` + +- **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. +- **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. +- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. +- **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five failed claims within one second, the backend stops calling the table for a second and fails requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). +- **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. + ## Upgrading across the v2 ingest envelope The NATS envelope changed shape in this release: the row now travels positionally, with `format`, `columns` and `row` replacing `data` — and the queue changed layout with it: boot deletes the earlier build's queue (below), so nothing an older version published reaches the new worker, which could not read it anyway (it carries no `format`, so there is no way to say which value belongs to which column). **Drain first** to keep what the old build had not yet inserted. diff --git a/go.mod b/go.mod index ca418e89..58b72f13 100644 --- a/go.mod +++ b/go.mod @@ -18,6 +18,11 @@ require ( github.com/ClickHouse/clickhouse-go/v2 v2.48.0 github.com/MicahParks/jwkset v0.11.3 github.com/MicahParks/keyfunc/v3 v3.8.2 + github.com/aws/aws-sdk-go-v2 v1.47.1 + github.com/aws/aws-sdk-go-v2/config v1.33.6 + github.com/aws/aws-sdk-go-v2/credentials v1.20.6 + github.com/aws/aws-sdk-go-v2/service/dynamodb v1.69.1 + github.com/aws/smithy-go v1.28.1 github.com/cockroachdb/pebble v1.1.5 github.com/dgraph-io/ristretto/v2 v2.4.2 github.com/dustin/go-humanize v1.0.1 @@ -72,6 +77,17 @@ require ( github.com/andybalholm/brotli v1.2.2 // indirect github.com/antithesishq/antithesis-sdk-go v0.7.2-default-no-op // indirect github.com/aws/aws-sdk-go v1.49.4 // indirect + github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.1 // indirect + github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.4 // indirect + github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.4 // indirect + github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.4 // indirect + github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 // indirect + github.com/aws/aws-sdk-go-v2/service/internal/endpoint-discovery v1.13.4 // indirect + github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.4 // indirect + github.com/aws/aws-sdk-go-v2/service/signin v1.10.1 // indirect + github.com/aws/aws-sdk-go-v2/service/sso v1.38.1 // indirect + github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.1 // indirect + github.com/aws/aws-sdk-go-v2/service/sts v1.51.1 // indirect github.com/aymanbagabas/go-osc52/v2 v2.0.1 // indirect github.com/beorn7/perks v1.0.1 // indirect github.com/bitfield/gotestdox v0.2.2 // indirect diff --git a/go.sum b/go.sum index 71dc2727..04ac3aaf 100644 --- a/go.sum +++ b/go.sum @@ -45,6 +45,38 @@ github.com/antithesishq/antithesis-sdk-go v0.7.2-default-no-op h1:p2zFsAzvhIpFya github.com/antithesishq/antithesis-sdk-go v0.7.2-default-no-op/go.mod h1:FQyySiasQQM8735Ddel3MRojmy4dA1IqCeyJ5jmPMbI= github.com/aws/aws-sdk-go v1.49.4 h1:qiXsqEeLLhdLgUIyfr5ot+N/dGPWALmtM1SetRmbUlY= github.com/aws/aws-sdk-go v1.49.4/go.mod h1:LF8svs817+Nz+DmiMQKTO3ubZ/6IaTpq3TjupRn3Eqk= +github.com/aws/aws-sdk-go-v2 v1.47.1 h1:uOIZnp4PK3ZhKI0dNrJrhTEsLxbpXHTAJlwoS1pvAtw= +github.com/aws/aws-sdk-go-v2 v1.47.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU= +github.com/aws/aws-sdk-go-v2/config v1.33.6 h1:MBjkSTLczek/UgiK+EYPIoRTqE7gP8vtW3OFbFo7Nug= +github.com/aws/aws-sdk-go-v2/config v1.33.6/go.mod h1:grRAFzdAZJrwcbasJRg2MPvIrVjtlfXllHssN6+E1JE= +github.com/aws/aws-sdk-go-v2/credentials v1.20.6 h1:NpAFXCU7NzXNkdGK3zQTtsRJ+3v9tZQV0xcdRw8uBdw= +github.com/aws/aws-sdk-go-v2/credentials v1.20.6/go.mod h1:mcZCoiPnyMvP8VMNbygNX5lLqSlkYJIMPODylQMurOk= +github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.1 h1:8gALAAmacnIXh+z6VkdDanv4/IkG5APdg4DZLDTmLog= +github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.1/go.mod h1:Z7IJhJU+poOdJjUR2wpyY21ossQ1XS/R3Lk9Msq5kM4= +github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.4 h1:CLq4+8UHCI+ZZYl/EuJxXovaIVN2xeeT8JV+dsApQ5E= +github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.4/go.mod h1:Wv4q5sAM04xAMkoOedxLx2inVf6K5FdxYp+A61L+q/0= +github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.4 h1:dD4MR81I7YkpEBRk6UP9rocC2QnT3qVuXwzlYTtfGEs= +github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.4/go.mod h1:EcXV1kAFd5XwSkDHlj94gnF3q5CkJyYiIJfH8N0VmrE= +github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.4 h1:7Wo47d/xn/7KttCSBd8EGYeZ7ULRFRkUHr6vkZPBzVQ= +github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.4/go.mod h1:tDB2IVC1xC3vX8o+6uRlzhTxP3g1b77CZXFX/oD2FnQ= +github.com/aws/aws-sdk-go-v2/service/dynamodb v1.69.1 h1:bKwiQA6SKqFXBO+1IwP/hTwCU5RlqeitG4gVvSuMN8U= +github.com/aws/aws-sdk-go-v2/service/dynamodb v1.69.1/go.mod h1:Gm+i2GlUsFNlzoBq8VXF44XHbKANn3tV8nYBBp3rN8Q= +github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 h1:bAdDl/HkGCcGPoe25ToSHEw23VIxt6CT5fLcg111BKg= +github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19/go.mod h1:KaUzbLxv4CeSxh6ZCl9B4m7CuFenS8kUEaDs+f/DQr4= +github.com/aws/aws-sdk-go-v2/service/internal/endpoint-discovery v1.13.4 h1:6HvmOQ1rBRrZ4qPJSWxd5szPKUsngXCwSw+V3UaJHmw= +github.com/aws/aws-sdk-go-v2/service/internal/endpoint-discovery v1.13.4/go.mod h1:zv2N29aiQUhG2XZNM9zgwCnAyVBdTBbcIpfNAlNmA20= +github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.4 h1:29SvnfGhXjTl8ONxFwbj2rs6lbhiFXD2CgFQmbT/bXY= +github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.4/go.mod h1:wm04I5DMuNVvZHFe/dHnUxincvNbbK7AiNBbYsQivek= +github.com/aws/aws-sdk-go-v2/service/signin v1.10.1 h1:DzCCWLzcIRQ77F3DEUljud7bEjTgFOIKXP52NmVRyhU= +github.com/aws/aws-sdk-go-v2/service/signin v1.10.1/go.mod h1:xpo/geVldu8payT375WekctUzopG/hBU7miiqItMUlw= +github.com/aws/aws-sdk-go-v2/service/sso v1.38.1 h1:Umtl/0YZhng4xndfW3lKJrYYP7NLEjI6bGXVomwLcs0= +github.com/aws/aws-sdk-go-v2/service/sso v1.38.1/go.mod h1:rRD/dnm7q0HYE/I5TMaPgkWyyUGLcwuxHLABsLnQ3e0= +github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.1 h1:orIWdNiLgzrhu/11RcPPKO/SBzUUymbUQuZbSPImghg= +github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.1/go.mod h1:skwM/xsbR/1ReUTesv9BhpJp1VjajR7DWQnuVLwiXsQ= +github.com/aws/aws-sdk-go-v2/service/sts v1.51.1 h1:0HOqZXRvMytH6bFHVIc0oJX07sZjfhz0zXtjs6gdE8s= +github.com/aws/aws-sdk-go-v2/service/sts v1.51.1/go.mod h1:26zA0GhDrLo+yiLI2yXWxqB1PdsShfLikoI7GOEgugM= +github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ= +github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc= github.com/aymanbagabas/go-osc52/v2 v2.0.1 h1:HwpRHbFMcZLEVr42D4p7XBqjyuxQH5SMiErDT4WkJ2k= github.com/aymanbagabas/go-osc52/v2 v2.0.1/go.mod h1:uYgXzlJ7ZpABp8OJ+exZzJJhRNQ2ASbcXHWsFqH8hp8= github.com/aymanbagabas/go-udiff v0.3.1 h1:LV+qyBQ2pqe0u42ZsUEtPiCaUoqgA9gYRDs3vj1nolY= diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go new file mode 100644 index 00000000..2ce30e6e --- /dev/null +++ b/internal/dedupe/dynamodb.go @@ -0,0 +1,605 @@ +package dedupe + +import ( + "context" + "crypto/rand" + "errors" + "fmt" + "log/slog" + "strconv" + "sync" + "time" + + "github.com/aws/aws-sdk-go-v2/aws" + "github.com/aws/aws-sdk-go-v2/aws/retry" + "github.com/aws/aws-sdk-go-v2/config" + "github.com/aws/aws-sdk-go-v2/service/dynamodb" + "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" + "github.com/aws/smithy-go" + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/metric" + "golang.org/x/sync/errgroup" + + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// The table's attributes. pk is the key from key.go and the only key +// attribute; ex is the table's TTL attribute. +const ( + attrKey = "pk" + attrState = "st" + attrExpiry = "ex" + attrToken = "tk" + + statePending = "1" + stateCommitted = "2" + + // A claim is live while now < ex; one whose ex has passed is absent + // to Reserve, whether or not TTL has deleted it yet. + condReserve = "attribute_not_exists(pk) OR ex <= :now" + condRelease = "tk = :tk AND st = :pending" + + // batchWriteMax is BatchWriteItem's per-call item limit. + batchWriteMax = 25 + // commitRounds bounds the BatchWriteItem rounds one chunk gets before + // its still-unprocessed items fail the Commit. + commitRounds = 8 + tokenBytes = 16 + + // opReserve is the operation the breaker watches: Release and Commit + // answers say nothing about whether a new Reserve would get through. + opReserve = "put_item" +) + +// DynamoConfig is the DynamoDB backend's wiring. Credentials are never here: +// the SDK's default chain finds them (EKS Pod Identity or IRSA in a pod, the +// environment or a profile locally). +type DynamoConfig struct { + // Table is the shared table every tenant's keys live in. Required. + Table string + // Region overrides the SDK chain's region (AWS_REGION) when set. + Region string + // Endpoint points the client at dynamodb-local. Tests and development + // only; it is also what unlocks CreateTable. + Endpoint string + // Timeout bounds each DynamoDB call, its SDK retries included. + // 0 = 250ms. + Timeout time.Duration + // MaxAttempts is the SDK retryer's attempts per call. 0 = 3. + MaxAttempts int + // RetryMode is "standard" (default) or "adaptive", which also rate-limits + // the client after throttles. + RetryMode string + // ReserveConcurrency bounds the parallel calls one Reserve, Commit or + // Release makes. 0 = 64. + ReserveConcurrency int +} + +func (c DynamoConfig) withDefaults() DynamoConfig { + if c.Timeout <= 0 { + c.Timeout = 250 * time.Millisecond + } + if c.MaxAttempts <= 0 { + c.MaxAttempts = 3 + } + if c.RetryMode == "" { + c.RetryMode = "standard" + } + if c.ReserveConcurrency <= 0 { + c.ReserveConcurrency = 64 + } + return c +} + +// dynamoAPI is the part of *dynamodb.Client the backend calls, so a unit test +// can inject throttles and unprocessed items. +type dynamoAPI interface { + PutItem(context.Context, *dynamodb.PutItemInput, ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) + BatchWriteItem(context.Context, *dynamodb.BatchWriteItemInput, ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) + DeleteItem(context.Context, *dynamodb.DeleteItemInput, ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) + DescribeTable(context.Context, *dynamodb.DescribeTableInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTableOutput, error) + DescribeTimeToLive(context.Context, *dynamodb.DescribeTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTimeToLiveOutput, error) + CreateTable(context.Context, *dynamodb.CreateTableInput, ...func(*dynamodb.Options)) (*dynamodb.CreateTableOutput, error) + UpdateTimeToLive(context.Context, *dynamodb.UpdateTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.UpdateTimeToLiveOutput, error) +} + +// Dynamo is the DynamoDB implementation: every tenant's keys in one shared +// table, so pods sharing the table share seen ids and Reserve's conditional +// write is atomic across all of them. WaveHouse never creates the table in +// production; CreateTable is for dynamodb-local. +type Dynamo struct { + api dynamoAPI + cfg DynamoConfig + now func() time.Time + breaker *breaker + metrics dynamoMetrics + // commitBackoff is the wait before retrying the attempt'th round of + // unprocessed items. + commitBackoff func(attempt int) time.Duration +} + +// NewDynamo builds the backend over a client from the SDK's default config +// chain. extra is appended to the chain's options (a test's static +// credentials, say). It dials nothing: Check does. +func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.LoadOptions) error) (*Dynamo, error) { + if cfg.Table == "" { + return nil, errors.New("dedupe: dynamodb table is required") + } + cfg = cfg.withDefaults() + retryer, err := newRetryer(cfg) + if err != nil { + return nil, err + } + opts := []func(*config.LoadOptions) error{config.WithRetryer(retryer)} + if cfg.Region != "" { + opts = append(opts, config.WithRegion(cfg.Region)) + } + awsCfg, err := config.LoadDefaultConfig(ctx, append(opts, extra...)...) + if err != nil { + return nil, fmt.Errorf("dedupe: aws config: %w", err) + } + client := dynamodb.NewFromConfig(awsCfg, func(o *dynamodb.Options) { + if cfg.Endpoint != "" { + o.BaseEndpoint = aws.String(cfg.Endpoint) + } + }) + return newDynamo(client, cfg), nil +} + +func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { + standard := func(o *retry.StandardOptions) { + o.MaxAttempts = cfg.MaxAttempts + o.MaxBackoff = 200 * time.Millisecond + } + switch cfg.RetryMode { + case "standard": + return func() aws.Retryer { return retry.NewStandard(standard) }, nil + case "adaptive": + return func() aws.Retryer { + return retry.NewAdaptiveMode(func(o *retry.AdaptiveModeOptions) { + o.StandardOptions = append(o.StandardOptions, standard) + }) + }, nil + } + return nil, fmt.Errorf("dedupe: dynamodb retry_mode %q: want standard or adaptive", cfg.RetryMode) +} + +func newDynamo(api dynamoAPI, cfg DynamoConfig) *Dynamo { + now := time.Now + return &Dynamo{ + api: api, + cfg: cfg.withDefaults(), + now: now, + breaker: newBreaker(now), + metrics: newDynamoMetrics(), + commitBackoff: func(attempt int) time.Duration { + return min(25*time.Millisecond< 0 { + ex = &types.AttributeValueMemberN{Value: strconv.FormatInt(expiresAt(s.d.now(), retention), 10)} + } + // BatchWriteItem refuses a key twice in one call; a caller merging + // claims from two Reserves could hand one over twice. + seen := make(map[string]bool, len(claims)) + writes := make([]types.WriteRequest, 0, len(claims)) + for _, c := range claims { + pk := AppendKey(nil, s.prefix, c.Key) + if seen[string(pk)] { + continue + } + seen[string(pk)] = true + item := map[string]types.AttributeValue{ + attrKey: &types.AttributeValueMemberB{Value: pk}, + attrState: &types.AttributeValueMemberN{Value: stateCommitted}, + attrToken: &types.AttributeValueMemberB{Value: []byte(c.Token)}, + } + if ex != nil { + item[attrExpiry] = ex + } + writes = append(writes, types.WriteRequest{PutRequest: &types.PutRequest{Item: item}}) + } + g, gctx := errgroup.WithContext(ctx) + g.SetLimit(s.d.cfg.ReserveConcurrency) + for start := 0; start < len(writes); start += batchWriteMax { + chunk := writes[start:min(start+batchWriteMax, len(writes))] + g.Go(func() error { return s.commitChunk(gctx, chunk) }) + } + return g.Wait() +} + +func (s *dynamoStore) commitChunk(ctx context.Context, writes []types.WriteRequest) error { + for attempt := 0; ; attempt++ { + var unprocessed []types.WriteRequest + err := s.d.call(ctx, "batch_write_item", func(ctx context.Context) error { + out, err := s.d.api.BatchWriteItem(ctx, &dynamodb.BatchWriteItemInput{ + RequestItems: map[string][]types.WriteRequest{s.d.cfg.Table: writes}, + }) + if err == nil { + unprocessed = out.UnprocessedItems[s.d.cfg.Table] + } + return err + }) + if err != nil { + return err + } + if len(unprocessed) == 0 { + return nil + } + if attempt+1 >= commitRounds { + return fmt.Errorf("%w: dynamodb batch_write_item: %d items still unprocessed", ErrUnavailable, len(unprocessed)) + } + s.d.metrics.unprocessed.Add(ctx, int64(len(unprocessed))) + writes = unprocessed + select { + case <-ctx.Done(): + return ctx.Err() + case <-time.After(s.d.commitBackoff(attempt)): + } + } +} + +// Release deletes each claim's item only while it is still that claim's +// pending item; a failed condition means the key lapsed, was re-claimed or +// was committed, and is left alone. +func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { + if len(claims) == 0 { + return nil + } + g, gctx := errgroup.WithContext(ctx) + g.SetLimit(s.d.cfg.ReserveConcurrency) + for _, c := range claims { + g.Go(func() error { + err := s.d.call(gctx, "delete_item", func(ctx context.Context) error { + _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ + TableName: &s.d.cfg.Table, + Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, + ConditionExpression: aws.String(condRelease), + ExpressionAttributeValues: map[string]types.AttributeValue{ + ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, + ":pending": &types.AttributeValueMemberN{Value: statePending}, + }, + }) + return err + }) + var gone *types.ConditionalCheckFailedException + if errors.As(err, &gone) { + return nil + } + return err + }) + } + return g.Wait() +} + +// Close is a no-op: the client is the Dynamo's, shared by every tenant. +func (s *dynamoStore) Close() error { return nil } + +// expiresAt is t+d in epoch seconds rounded up, so a claim or commit never +// ends before it was asked to: TTL attributes are whole seconds. +func expiresAt(t time.Time, d time.Duration) int64 { + end := t.Add(d) + sec := end.Unix() + if end.Nanosecond() > 0 { + sec++ + } + return sec +} + +func newToken() string { + b := make([]byte, tokenBytes) + _, _ = rand.Read(b) // crypto/rand.Read never fails + return string(b) +} + +// classify maps a DynamoDB error onto the contract: a condition failure is +// returned as is for the caller to read, anything retrying later can cure +// wraps ErrUnavailable (503), and the rest — a missing table, denied access, +// a malformed request — is a configuration bug (500). +func classify(op string, err error) error { + if err == nil { + return nil + } + var cond *types.ConditionalCheckFailedException + if errors.As(err, &cond) { + return err + } + if transient(err) { + return fmt.Errorf("%w: dynamodb %s: %w", ErrUnavailable, op, err) + } + return fmt.Errorf("dynamodb %s: %w", op, err) +} + +func transient(err error) bool { + if errors.Is(err, context.DeadlineExceeded) { + return true + } + if (retry.RetryableConnectionError{}).IsErrorRetryable(err) == aws.TrueTernary { + return true + } + var api smithy.APIError + if !errors.As(err, &api) { + return false + } + code := api.ErrorCode() + if _, ok := retry.DefaultThrottleErrorCodes[code]; ok { + return true + } + if _, ok := retry.DefaultRetryableErrorCodes[code]; ok { + return true + } + switch code { + case "InternalServerError", "ServiceUnavailable", "ReplicatedWriteConflictException": + return true + } + return api.ErrorFault() == smithy.FaultServer +} + +// breaker short-circuits Reserve for a second after breakerTrips consecutive +// unavailable answers inside a second, so a throttled or unreachable table +// fails requests fast instead of spending every one's full timeout. +type breaker struct { + mu sync.Mutex + now func() time.Time + fails int + since time.Time + openUntil time.Time +} + +const ( + breakerTrips = 5 + breakerWindow = time.Second + breakerCool = time.Second +) + +var errBreakerOpen = fmt.Errorf("%w: dynamodb is failing; short-circuited", ErrUnavailable) + +func newBreaker(now func() time.Time) *breaker { return &breaker{now: now} } + +func (b *breaker) allow() error { + b.mu.Lock() + defer b.mu.Unlock() + if b.now().Before(b.openUntil) { + return errBreakerOpen + } + return nil +} + +func (b *breaker) record(err error) { + b.mu.Lock() + defer b.mu.Unlock() + if !errors.Is(err, ErrUnavailable) { + b.fails = 0 + return + } + now := b.now() + if b.fails == 0 || now.Sub(b.since) > breakerWindow { + b.fails, b.since = 0, now + } + b.fails++ + if b.fails >= breakerTrips { + b.fails = 0 + b.openUntil = now.Add(breakerCool) + } +} + +type dynamoMetrics struct { + requests metric.Int64Counter + duration metric.Float64Histogram + unprocessed metric.Int64Counter + shorted metric.Int64Counter +} + +func newDynamoMetrics() dynamoMetrics { + meter := otel.Meter("wavehouse-dedupe") + requests, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_requests_total", + metric.WithDescription("DynamoDB dedupe requests by operation and outcome (ok, condition_failed, unavailable, error)")) + duration, _ := meter.Float64Histogram("wavehouse_dedupe_dynamodb_request_duration_seconds", + metric.WithDescription("DynamoDB dedupe request latency, SDK retries included"), metric.WithUnit("s")) + unprocessed, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_unprocessed_items_total", + metric.WithDescription("Commit items DynamoDB left unprocessed and the backend retried")) + shorted, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_short_circuits_total", + metric.WithDescription("Reserves refused without a request while DynamoDB was failing")) + return dynamoMetrics{requests: requests, duration: duration, unprocessed: unprocessed, shorted: shorted} +} + +func (m dynamoMetrics) record(ctx context.Context, op string, took time.Duration, err error) { + outcome := "ok" + var cond *types.ConditionalCheckFailedException + switch { + case err == nil: + case errors.As(err, &cond): + outcome = "condition_failed" + case errors.Is(err, ErrUnavailable): + outcome = "unavailable" + default: + outcome = "error" + } + ctx = context.WithoutCancel(ctx) + m.requests.Add(ctx, 1, metric.WithAttributes(attribute.String("op", op), attribute.String("outcome", outcome))) + m.duration.Record(ctx, took.Seconds(), metric.WithAttributes(attribute.String("op", op))) +} + +func (m dynamoMetrics) shortCircuit(ctx context.Context) { m.shorted.Add(ctx, 1) } diff --git a/internal/dedupe/dynamodb_bench_test.go b/internal/dedupe/dynamodb_bench_test.go new file mode 100644 index 00000000..ad7239ed --- /dev/null +++ b/internal/dedupe/dynamodb_bench_test.go @@ -0,0 +1,77 @@ +//go:build dynamobench + +// Manual latency benchmark for the DynamoDB backend, never run by CI. Point it +// at an existing table (the credentials and region come from the SDK chain): +// +// DEDUPE_BENCH_TABLE=wavehouse-dedupe-dev go test -tags dynamobench \ +// -run '^$' -bench Dynamo -benchtime 2000x ./internal/dedupe/ +// +// DEDUPE_BENCH_ENDPOINT=http://localhost:8000 runs it against dynamodb-local +// instead, creating the table there. +package dedupe + +import ( + "fmt" + "os" + "sync/atomic" + "testing" + "time" +) + +var benchSeq atomic.Uint64 + +func benchDynamo(b *testing.B) *Managed { + b.Helper() + cfg := DynamoConfig{Table: os.Getenv("DEDUPE_BENCH_TABLE"), Endpoint: os.Getenv("DEDUPE_BENCH_ENDPOINT")} + if cfg.Table == "" { + b.Skip("DEDUPE_BENCH_TABLE is not set") + } + if cfg.Endpoint != "" { + cfg.Timeout = 5 * time.Second // dynamodb-local is far slower than the service + } + d, err := NewDynamo(b.Context(), cfg) + if err != nil { + b.Fatal(err) + } + if cfg.Endpoint != "" { + if err := d.CreateTable(b.Context()); err != nil { + b.Fatal(err) + } + } + if err := d.Check(b.Context()); err != nil { + b.Fatal(err) + } + m := d.Tenant("bench") + if err := m.Apply(true); err != nil { + b.Fatal(err) + } + return m +} + +// benchKeys are n ids no run has used, with a short retention so the table +// forgets them. +func benchKeys(n int) []Key { + run := time.Now().UnixNano() + out := make([]Key, n) + for i := range out { + out[i] = Key{Table: "bench", ID: fmt.Sprintf("%d-%d", run, benchSeq.Add(1))} + } + return out +} + +func benchReserveCommit(b *testing.B, window int) { + m := benchDynamo(b) + b.ResetTimer() + for b.Loop() { + claims, err := m.Reserve(b.Context(), benchKeys(window), DefaultLease) + if err != nil { + b.Fatal(err) + } + if err := m.Commit(b.Context(), claims, time.Hour); err != nil { + b.Fatal(err) + } + } +} + +func BenchmarkDynamo_ReserveCommit1(b *testing.B) { benchReserveCommit(b, 1) } +func BenchmarkDynamo_ReserveCommit256(b *testing.B) { benchReserveCommit(b, 256) } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go new file mode 100644 index 00000000..8648c938 --- /dev/null +++ b/internal/dedupe/dynamodb_test.go @@ -0,0 +1,357 @@ +package dedupe + +import ( + "context" + "errors" + "fmt" + "net" + "sync" + "sync/atomic" + "testing" + "time" + + "github.com/aws/aws-sdk-go-v2/aws" + "github.com/aws/aws-sdk-go-v2/service/dynamodb" + "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" + "github.com/aws/smithy-go" + smithyhttp "github.com/aws/smithy-go/transport/http" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// fakeDynamo answers each operation through its func, or with success when +// that is nil. The DynamoDB semantics themselves are tested against +// dynamodb-local (tests/integration); this is for the error paths it cannot +// produce. +type fakeDynamo struct { + put func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) + batch func(*dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) + del func(*dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) + describe func() (*dynamodb.DescribeTableOutput, error) + ttl func() (*dynamodb.DescribeTimeToLiveOutput, error) +} + +func (f *fakeDynamo) PutItem(_ context.Context, in *dynamodb.PutItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) { + if f.put == nil { + return &dynamodb.PutItemOutput{}, nil + } + return f.put(in) +} + +func (f *fakeDynamo) BatchWriteItem(_ context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { + if f.batch == nil { + return &dynamodb.BatchWriteItemOutput{}, nil + } + return f.batch(in) +} + +func (f *fakeDynamo) DeleteItem(_ context.Context, in *dynamodb.DeleteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) { + if f.del == nil { + return &dynamodb.DeleteItemOutput{}, nil + } + return f.del(in) +} + +func (f *fakeDynamo) DescribeTable(context.Context, *dynamodb.DescribeTableInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTableOutput, error) { + return f.describe() +} + +func (f *fakeDynamo) DescribeTimeToLive(context.Context, *dynamodb.DescribeTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTimeToLiveOutput, error) { + return f.ttl() +} + +func (f *fakeDynamo) CreateTable(context.Context, *dynamodb.CreateTableInput, ...func(*dynamodb.Options)) (*dynamodb.CreateTableOutput, error) { + return nil, errors.New("not used") +} + +func (f *fakeDynamo) UpdateTimeToLive(context.Context, *dynamodb.UpdateTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.UpdateTimeToLiveOutput, error) { + return nil, errors.New("not used") +} + +func apiErr(code string, fault smithy.ErrorFault) error { + return &smithy.GenericAPIError{Code: code, Message: "injected", Fault: fault} +} + +func openFake(t *testing.T, f *fakeDynamo) (*Dynamo, Deduplicator) { + t.Helper() + d := newDynamo(f, DynamoConfig{Table: "dedupe"}) + d.commitBackoff = func(int) time.Duration { return 0 } + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + t.Cleanup(func() { _ = m.Close() }) + return d, m +} + +func keys(ids ...string) []Key { + out := make([]Key, len(ids)) + for i, id := range ids { + out[i] = Key{Table: "events", ID: id} + } + return out +} + +func TestClassify(t *testing.T) { + t.Parallel() + for _, tc := range []struct { + name string + err error + unavailable bool + }{ + {"throttled", apiErr("ThrottlingException", smithy.FaultClient), true}, + {"over provisioned throughput", &types.ProvisionedThroughputExceededException{}, true}, + {"account request limit", apiErr("RequestLimitExceeded", smithy.FaultClient), true}, + {"internal error", &types.InternalServerError{}, true}, + {"unknown server fault", apiErr("Whatever", smithy.FaultServer), true}, + {"request timeout", apiErr("RequestTimeoutException", smithy.FaultClient), true}, + {"multi-region write conflict", &types.ReplicatedWriteConflictException{}, true}, + {"deadline", fmt.Errorf("op: %w", context.DeadlineExceeded), true}, + {"connection refused", &smithyhttp.RequestSendError{Err: &net.OpError{Op: "dial", Err: errors.New("refused")}}, true}, + {"missing table", &types.ResourceNotFoundException{}, false}, + {"access denied", apiErr("AccessDeniedException", smithy.FaultClient), false}, + {"validation", apiErr("ValidationException", smithy.FaultClient), false}, + {"caller went away", context.Canceled, false}, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + err := classify("put_item", tc.err) + assert.ErrorIs(t, err, tc.err, "the cause stays reachable") + assert.Equal(t, tc.unavailable, errors.Is(err, ErrUnavailable)) + }) + } + assert.NoError(t, classify("put_item", nil)) + ccf := &types.ConditionalCheckFailedException{} + assert.Same(t, error(ccf), classify("put_item", ccf), "a condition failure is an answer, not an error") +} + +func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { + t.Parallel() + _, m := openFake(t, &fakeDynamo{put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + switch id[len(id)-1] { + case 'd': + return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: stateCommitted}}} + case 'f': + return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: statePending}}} + } + assert.Equal(t, condReserve, aws.ToString(in.ConditionExpression)) + assert.Equal(t, types.ReturnValuesOnConditionCheckFailureAllOld, in.ReturnValuesOnConditionCheckFailure) + return &dynamodb.PutItemOutput{}, nil + }}) + claims, err := m.Reserve(t.Context(), keys("new", "old", "inf"), time.Minute) + require.NoError(t, err) + assert.Equal(t, []Status{Claimed, Duplicate, InFlight}, []Status{claims[0].Status, claims[1].Status, claims[2].Status}) + assert.Len(t, claims[0].Token, tokenBytes) + assert.Empty(t, claims[1].Token) +} + +func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { + t.Parallel() + var mu sync.Mutex + putTokens := map[string]string{} + var released []string + _, m := openFake(t, &fakeDynamo{ + put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + mu.Lock() + putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) + mu.Unlock() + switch id[len(id)-3:] { + case "dup": + return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: stateCommitted}}} + case "bad": + return nil, &types.InternalServerError{} + } + return &dynamodb.PutItemOutput{}, nil + }, + del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + mu.Lock() + defer mu.Unlock() + assert.Equal(t, putTokens[id], string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value), "released by the token it was put with") + released = append(released, id[len(id)-3:]) + return &dynamodb.DeleteItemOutput{}, nil + }, + }) + _, err := m.Reserve(t.Context(), keys("ok1", "dup", "bad", "ok2"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + assert.ElementsMatch(t, []string{"ok1", "bad", "ok2"}, released, "the failed put may have landed; the duplicate was never ours") +} + +func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { + t.Parallel() + var calls atomic.Int64 + var mu sync.Mutex + written := map[string]int{} + heldBack := map[string]bool{} + _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + calls.Add(1) + reqs := in.RequestItems["dedupe"] + assert.LessOrEqual(t, len(reqs), batchWriteMax) + // Leave the last item of every call unprocessed once. + mu.Lock() + defer mu.Unlock() + var left []types.WriteRequest + for i, r := range reqs { + pk := string(r.PutRequest.Item[attrKey].(*types.AttributeValueMemberB).Value) + assert.Equal(t, stateCommitted, r.PutRequest.Item[attrState].(*types.AttributeValueMemberN).Value) + assert.Contains(t, r.PutRequest.Item, attrExpiry) + if i == len(reqs)-1 && !heldBack[pk] && len(reqs) > 1 { + heldBack[pk] = true + left = append(left, r) + continue + } + written[pk]++ + } + return &dynamodb.BatchWriteItemOutput{UnprocessedItems: map[string][]types.WriteRequest{"dedupe": left}}, nil + }}) + ids := make([]string, 60) + for i := range ids { + ids[i] = fmt.Sprint(i) + } + claims := make([]Claim, 0, len(ids)+1) + for _, k := range keys(ids...) { + claims = append(claims, Claim{Key: k, Status: Claimed, Token: "t"}) + } + claims = append(claims, claims[0]) + require.NoError(t, m.Commit(t.Context(), claims, time.Hour)) + assert.Len(t, written, 60, "a key handed over twice is written once") + for pk, n := range written { + assert.Equal(t, 1, n, "%q", pk) + } + assert.Equal(t, int64(6), calls.Load(), "3 chunks, each retried once") +} + +func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { + t.Parallel() + _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + return &dynamodb.BatchWriteItemOutput{UnprocessedItems: in.RequestItems}, nil + }}) + err := m.Commit(t.Context(), []Claim{{Key: keys("a")[0], Status: Claimed, Token: "t"}}, 0) + require.ErrorIs(t, err, ErrUnavailable) +} + +func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { + t.Parallel() + _, m := openFake(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + assert.Equal(t, condRelease, aws.ToString(in.ConditionExpression)) + id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + if id[len(id)-1] == 'x' { + return nil, &types.ResourceNotFoundException{} + } + return nil, &types.ConditionalCheckFailedException{} + }}) + claim := func(id string) []Claim { return []Claim{{Key: keys(id)[0], Status: Claimed, Token: "t"}} } + require.NoError(t, m.Release(t.Context(), claim("gone"))) + err := m.Release(t.Context(), claim("x")) + require.Error(t, err) + assert.False(t, errors.Is(err, ErrUnavailable)) +} + +func TestDynamo_BreakerShortCircuitsReserve(t *testing.T) { + t.Parallel() + var puts atomic.Int64 + var down atomic.Bool + down.Store(true) + d, m := openFake(t, &fakeDynamo{put: func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + puts.Add(1) + if down.Load() { + return nil, &types.ProvisionedThroughputExceededException{} + } + return &dynamodb.PutItemOutput{}, nil + }}) + now := time.Unix(1_000_000, 0) + var clock sync.Mutex + d.breaker.now = func() time.Time { clock.Lock(); defer clock.Unlock(); return now } + for range breakerTrips { + _, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + } + _, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, errBreakerOpen) + assert.Equal(t, int64(breakerTrips), puts.Load(), "the open breaker sent nothing") + + down.Store(false) + clock.Lock() + now = now.Add(breakerCool) + clock.Unlock() + c, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.NoError(t, err, "it closes after the cool-down") + assert.Equal(t, Claimed, c[0].Status) +} + +func TestBreaker_FailuresSpreadOutDoNotTrip(t *testing.T) { + t.Parallel() + now := time.Unix(1_000_000, 0) + b := newBreaker(func() time.Time { return now }) + fail := fmt.Errorf("%w: x", ErrUnavailable) + for range 3 * breakerTrips { + b.record(fail) + now = now.Add(breakerWindow/(breakerTrips-1) + time.Millisecond) + } + require.NoError(t, b.allow()) + now = now.Add(2 * breakerWindow) + for range breakerTrips - 1 { + b.record(fail) + } + b.record(nil) + b.record(fail) + require.NoError(t, b.allow(), "a success resets the count") +} + +func TestDynamo_Check(t *testing.T) { + t.Parallel() + good := &dynamodb.DescribeTableOutput{Table: &types.TableDescription{ + KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeB}}, + }} + ttlOn := &dynamodb.DescribeTimeToLiveOutput{TimeToLiveDescription: &types.TimeToLiveDescription{ + AttributeName: aws.String("ex"), TimeToLiveStatus: types.TimeToLiveStatusEnabled, + }} + check := func(table *dynamodb.DescribeTableOutput, ttl *dynamodb.DescribeTimeToLiveOutput, ttlErr error) error { + f := &fakeDynamo{ + describe: func() (*dynamodb.DescribeTableOutput, error) { return table, nil }, + ttl: func() (*dynamodb.DescribeTimeToLiveOutput, error) { return ttl, ttlErr }, + } + return newDynamo(f, DynamoConfig{Table: "dedupe"}).Check(t.Context()) + } + require.NoError(t, check(good, ttlOn, nil)) + require.NoError(t, check(good, &dynamodb.DescribeTimeToLiveOutput{}, nil), "no TTL is a warning") + require.ErrorIs(t, check(good, nil, &types.InternalServerError{}), ErrUnavailable) + + withRange := &dynamodb.DescribeTableOutput{Table: &types.TableDescription{ + KeySchema: []types.KeySchemaElement{ + {AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}, + {AttributeName: aws.String("sk"), KeyType: types.KeyTypeRange}, + }, + }} + assert.ErrorContains(t, check(withRange, ttlOn, nil), "key schema") +} + +func TestDynamo_Config(t *testing.T) { + t.Parallel() + c := DynamoConfig{}.withDefaults() + assert.Equal(t, DynamoConfig{Timeout: 250 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard", ReserveConcurrency: 64}, c) + for _, mode := range []string{"standard", "adaptive"} { + r, err := newRetryer(DynamoConfig{RetryMode: mode, MaxAttempts: 4}) + require.NoError(t, err) + assert.Equal(t, 4, r().MaxAttempts()) + } + _, err := newRetryer(DynamoConfig{RetryMode: "legacy"}) + require.Error(t, err) + + _, err = NewDynamo(t.Context(), DynamoConfig{}) + require.ErrorContains(t, err, "table is required") + _, err = NewDynamo(t.Context(), DynamoConfig{Table: "t", RetryMode: "legacy"}) + require.Error(t, err) + d, err := NewDynamo(t.Context(), DynamoConfig{Table: "t", Region: "us-east-1"}) + require.NoError(t, err) + require.ErrorIs(t, d.CreateTable(t.Context()), ErrCreateTableNeedsEndpoint, "never against real AWS") +} + +func TestExpiresAt(t *testing.T) { + t.Parallel() + base := time.Unix(100, 0) + assert.Equal(t, int64(101), expiresAt(base, time.Second)) + assert.Equal(t, int64(102), expiresAt(base, 1500*time.Millisecond), "rounded up: never ends early") + assert.Equal(t, int64(102), expiresAt(base.Add(time.Nanosecond), time.Second)) +} diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go new file mode 100644 index 00000000..192734a9 --- /dev/null +++ b/tests/integration/dedupe_dynamodb_test.go @@ -0,0 +1,302 @@ +//go:build integration + +package tests + +import ( + "context" + "errors" + "fmt" + "io" + "net/http" + "strconv" + "strings" + "sync" + "sync/atomic" + "testing" + "time" + + "github.com/aws/aws-sdk-go-v2/aws" + "github.com/aws/aws-sdk-go-v2/config" + "github.com/aws/aws-sdk-go-v2/credentials" + "github.com/aws/aws-sdk-go-v2/service/dynamodb" + "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/dedupe/dedupetest" +) + +var dynamoTables atomic.Uint64 + +// newDynamoTable names a fresh table on dynamodb-local for one test. +func newDynamoTable() string { + return fmt.Sprintf("dedupe_%d", dynamoTables.Add(1)) +} + +// dynamoClient is one client — one pod's view — over table on +// dynamodb-local, through the production constructor. +func dynamoClient(t *testing.T, table string, cfg dedupe.DynamoConfig, extra ...func(*config.LoadOptions) error) *dedupe.Dynamo { + t.Helper() + cfg.Table, cfg.Endpoint, cfg.Region = table, env(t).dynamoEndpoint, "us-east-1" + if cfg.Timeout == 0 { + // dynamodb-local under a parallel suite is slower than the real thing. + cfg.Timeout = 5 * time.Second + } + opts := append([]func(*config.LoadOptions) error{ + config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider("local", "local", "")), + }, extra...) + d, err := dedupe.NewDynamo(t.Context(), cfg, opts...) + require.NoError(t, err) + return d +} + +// rawDynamo is a plain client, for reading and planting items directly. +func rawDynamo(t *testing.T) *dynamodb.Client { + t.Helper() + return dynamodb.New(dynamodb.Options{ + Region: "us-east-1", + BaseEndpoint: aws.String(env(t).dynamoEndpoint), + Credentials: credentials.NewStaticCredentialsProvider("local", "local", ""), + }) +} + +// faultyHTTP answers matching requests itself instead of sending them. +type faultyHTTP struct { + next *http.Client + fault func(target string) (*http.Response, error, bool) +} + +func (f *faultyHTTP) Do(r *http.Request) (*http.Response, error) { + if resp, err, ok := f.fault(r.Header.Get("X-Amz-Target")); ok { + return resp, err + } + return f.next.Do(r) +} + +// awsError is a DynamoDB JSON error response. +func awsError(status int, code string) *http.Response { + body := fmt.Sprintf(`{"__type":"com.amazonaws.dynamodb.v20120810#%s","message":"injected"}`, code) + return &http.Response{ + StatusCode: status, + Header: http.Header{"Content-Type": {"application/x-amz-json-1.0"}}, + Body: io.NopCloser(strings.NewReader(body)), + } +} + +const putItem = "DynamoDB_20120810.PutItem" + +func TestDedupeDynamo_Conformance(t *testing.T) { + t.Parallel() + dedupetest.Run(t, func(t *testing.T) dedupetest.Harness { + table := newDynamoTable() + // failAfter < 0 is off; otherwise the put after that many fails once. + var failAfter, puts atomic.Int64 + failAfter.Store(-1) + fault := config.WithHTTPClient(&faultyHTTP{next: http.DefaultClient, fault: func(target string) (*http.Response, error, bool) { + if target != putItem || failAfter.Load() < 0 || puts.Add(1) <= failAfter.Load() { + return nil, nil, false + } + failAfter.Store(-1) + return awsError(http.StatusBadRequest, "ValidationException"), nil, true + }}) + d := dynamoClient(t, table, dedupe.DynamoConfig{}, fault) + require.NoError(t, d.CreateTable(t.Context())) + require.NoError(t, d.Check(t.Context())) + return dedupetest.Harness{ + Factory: d.Tenant, + Peer: dynamoClient(t, table, dedupe.DynamoConfig{}).Tenant, + FailNextReserve: func(n int) { + puts.Store(0) + failAfter.Store(int64(n)) + }, + } + }) +} + +// 32 clients — 32 pods — race one id: DynamoDB's condition, not anything in +// process, is what lets exactly one through. +func TestDedupeDynamo_ThirtyTwoClientsOneID(t *testing.T) { + t.Parallel() + table := newDynamoTable() + first := dynamoClient(t, table, dedupe.DynamoConfig{}) + require.NoError(t, first.CreateTable(t.Context())) + const n = 32 + stores := make([]*dedupe.Managed, n) + for i := range stores { + stores[i] = dynamoClient(t, table, dedupe.DynamoConfig{}).Tenant("acme") + require.NoError(t, stores[i].Apply(true)) + } + k := []dedupe.Key{{Table: "events", ID: "e1"}} + race := func() map[dedupe.Status][]dedupe.Claim { + got := make([]dedupe.Claim, n) + start := make(chan struct{}) + var wg sync.WaitGroup + for i, s := range stores { + wg.Go(func() { + <-start + c, err := s.Reserve(context.Background(), k, time.Minute) + if assert.NoError(t, err) { + got[i] = c[0] + } + }) + } + close(start) + wg.Wait() + by := map[dedupe.Status][]dedupe.Claim{} + for _, c := range got { + by[c.Status] = append(by[c.Status], c) + } + return by + } + by := race() + require.Len(t, by[dedupe.Claimed], 1, "exactly one client claims the id") + assert.Len(t, by[dedupe.InFlight], n-1) + require.NoError(t, stores[0].Commit(t.Context(), by[dedupe.Claimed], 0)) + assert.Len(t, race()[dedupe.Duplicate], n, "and every client then sees it committed") +} + +func TestDedupeDynamo_Throttled(t *testing.T) { + t.Parallel() + table := newDynamoTable() + require.NoError(t, dynamoClient(t, table, dedupe.DynamoConfig{}).CreateTable(t.Context())) + for _, code := range []string{"ThrottlingException", "ProvisionedThroughputExceededException", "RequestLimitExceeded"} { + t.Run(code, func(t *testing.T) { + t.Parallel() + var sent atomic.Int64 + d := dynamoClient(t, table, dedupe.DynamoConfig{MaxAttempts: 2}, config.WithHTTPClient(&faultyHTTP{ + next: http.DefaultClient, + fault: func(target string) (*http.Response, error, bool) { + if target != putItem { + return nil, nil, false + } + sent.Add(1) + return awsError(http.StatusBadRequest, code), nil, true + }, + })) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + _, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "e1"}}, time.Minute) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "a throttle is worth retrying: 503") + assert.Equal(t, int64(2), sent.Load(), "the SDK retried it once first") + }) + } +} + +func TestDedupeDynamo_Unreachable(t *testing.T) { + t.Parallel() + d, err := dedupe.NewDynamo(t.Context(), dedupe.DynamoConfig{ + Table: "dedupe", Region: "us-east-1", Endpoint: "http://127.0.0.1:1", MaxAttempts: 1, + }, config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider("local", "local", ""))) + require.NoError(t, err) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + k := []dedupe.Key{{Table: "events", ID: "e1"}} + for range 5 { + _, err = m.Reserve(t.Context(), k, time.Minute) + require.ErrorIs(t, err, dedupe.ErrUnavailable) + } + _, err = m.Reserve(t.Context(), k, time.Minute) + require.ErrorIs(t, err, dedupe.ErrUnavailable) + assert.Contains(t, err.Error(), "short-circuited", "five failures in a second open the breaker") + assert.ErrorIs(t, d.Check(t.Context()), dedupe.ErrUnavailable) +} + +func TestDedupeDynamo_ConfigErrorsAreNotUnavailable(t *testing.T) { + t.Parallel() + d := dynamoClient(t, "no_such_table", dedupe.DynamoConfig{}) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + _, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "e1"}}, time.Minute) + require.Error(t, err) + var missing *types.ResourceNotFoundException + assert.ErrorAs(t, err, &missing) + assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a missing table is a config bug: 500, not 503") + require.Error(t, d.Check(t.Context())) +} + +func TestDedupeDynamo_Check(t *testing.T) { + t.Parallel() + raw := rawDynamo(t) + table := newDynamoTable() + _, err := raw.CreateTable(t.Context(), &dynamodb.CreateTableInput{ + TableName: aws.String(table), + BillingMode: types.BillingModePayPerRequest, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeS}}, + KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, + }) + require.NoError(t, err) + assert.ErrorContains(t, dynamoClient(t, table, dedupe.DynamoConfig{}).Check(t.Context()), "must be binary") + + fresh := newDynamoTable() + d := dynamoClient(t, fresh, dedupe.DynamoConfig{}) + require.NoError(t, d.CreateTable(t.Context())) + require.NoError(t, d.CreateTable(t.Context()), "an existing table is left alone") + ttl, err := raw.DescribeTimeToLive(t.Context(), &dynamodb.DescribeTimeToLiveInput{TableName: aws.String(fresh)}) + require.NoError(t, err) + assert.Equal(t, "ex", aws.ToString(ttl.TimeToLiveDescription.AttributeName)) + assert.Equal(t, types.TimeToLiveStatusEnabled, ttl.TimeToLiveDescription.TimeToLiveStatus) +} + +// Expiry is the item's ex, in epoch seconds, and never depends on TTL having +// deleted the item. +func TestDedupeDynamo_Expiry(t *testing.T) { + t.Parallel() + raw := rawDynamo(t) + table := newDynamoTable() + d := dynamoClient(t, table, dedupe.DynamoConfig{}) + require.NoError(t, d.CreateTable(t.Context())) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + pk := func(id string) []byte { + return dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "events", ID: id}) + } + item := func(id string) map[string]types.AttributeValue { + out, err := raw.GetItem(t.Context(), &dynamodb.GetItemInput{ + TableName: aws.String(table), ConsistentRead: aws.Bool(true), + Key: map[string]types.AttributeValue{"pk": &types.AttributeValueMemberB{Value: pk(id)}}, + }) + require.NoError(t, err) + return out.Item + } + num := func(av types.AttributeValue) int64 { + n, err := strconv.ParseInt(av.(*types.AttributeValueMemberN).Value, 10, 64) + require.NoError(t, err) + return n + } + + before := time.Now() + claims, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "kept"}, {Table: "events", ID: "brief"}}, 30*time.Second) + require.NoError(t, err) + pending := item("kept") + assert.Equal(t, "1", pending["st"].(*types.AttributeValueMemberN).Value) + assert.InDelta(t, before.Add(30*time.Second).Unix(), num(pending["ex"]), 2, "a pending item's ex is its lease end") + + require.NoError(t, m.Commit(t.Context(), claims[:1], 0)) + require.NoError(t, m.Commit(t.Context(), claims[1:], time.Hour)) + assert.NotContains(t, item("kept"), "ex", "retention 0 writes no ex, so TTL never takes it") + assert.InDelta(t, before.Add(time.Hour).Unix(), num(item("brief")["ex"]), 2, "a commit's ex is its retention end") + + // TTL deletes lazily; an item whose ex has passed is absent all the same. + for _, st := range []string{"1", "2"} { + _, err = raw.PutItem(t.Context(), &dynamodb.PutItemInput{TableName: aws.String(table), Item: map[string]types.AttributeValue{ + "pk": &types.AttributeValueMemberB{Value: pk("stale-" + st)}, + "st": &types.AttributeValueMemberN{Value: st}, + "ex": &types.AttributeValueMemberN{Value: strconv.FormatInt(time.Now().Add(-time.Minute).Unix(), 10)}, + "tk": &types.AttributeValueMemberB{Value: []byte("old")}, + }}) + require.NoError(t, err) + } + got, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "stale-1"}, {Table: "events", ID: "stale-2"}, {Table: "events", ID: "kept"}}, time.Minute) + require.NoError(t, err) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed, dedupe.Duplicate}, + []dedupe.Status{got[0].Status, got[1].Status, got[2].Status}) +} + +func TestDedupeDynamo_CreateTableNeedsEndpoint(t *testing.T) { + t.Parallel() + d, err := dedupe.NewDynamo(t.Context(), dedupe.DynamoConfig{Table: "dedupe", Region: "us-east-1"}, + config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider("local", "local", ""))) + require.NoError(t, err) + require.ErrorIs(t, d.CreateTable(t.Context()), dedupe.ErrCreateTableNeedsEndpoint) +} diff --git a/tests/integration/setup_test.go b/tests/integration/setup_test.go index a01064a2..9d166690 100644 --- a/tests/integration/setup_test.go +++ b/tests/integration/setup_test.go @@ -51,6 +51,9 @@ type testEnv struct { embeddedMQ mq.Broker baseURL string // the wired API server, e.g. http://127.0.0.1:41234 registry *discovery.SchemaRegistry + // dynamoEndpoint is dynamodb-local, for the DynamoDB dedupe backend's + // tests; the wired app does not use it. + dynamoEndpoint string } var sharedEnv *testEnv @@ -137,6 +140,13 @@ func setup() (int, func()) { _ = ch.container.Terminate(context.Background()) }) + ddb, endpoint, err := startDynamoDBLocal(ctx) + if err != nil { + fmt.Fprintf(os.Stderr, "integration setup: dynamodb-local: %v\n", err) + return 1, cleanup + } + cleanups.push(func() { _ = ddb.Terminate(context.Background()) }) + settingsDir, err := writeTestSettings(ch) if err != nil { fmt.Fprintf(os.Stderr, "integration setup: settings: %v\n", err) @@ -199,6 +209,8 @@ func setup() (int, func()) { embeddedMQ: a.MQ(), baseURL: baseURL, registry: a.Registry(), + + dynamoEndpoint: endpoint, } return 0, cleanup } @@ -383,6 +395,28 @@ func startClickHouse(ctx context.Context) (*chInstance, error) { return ch, nil } +// startDynamoDBLocal starts dynamodb-local in memory (no volume) and +// returns it with its endpoint URL. +func startDynamoDBLocal(ctx context.Context) (testcontainers.Container, string, error) { + container, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ + ContainerRequest: testcontainers.ContainerRequest{ + Image: "amazon/dynamodb-local:3.3.1", + Cmd: []string{"-jar", "DynamoDBLocal.jar", "-inMemory"}, + ExposedPorts: []string{"8000/tcp"}, + WaitingFor: wait.ForListeningPort("8000/tcp").WithStartupTimeout(60 * time.Second), + }, + Started: true, + }) + if err != nil { + return nil, "", fmt.Errorf("start container: %w", err) + } + endpoint, err := container.PortEndpoint(ctx, "8000/tcp", "http") + if err != nil { + return container, "", fmt.Errorf("endpoint: %w", err) + } + return container, endpoint, nil +} + func waitForNativeReady(ctx context.Context, conn driver.Conn, timeout time.Duration) error { pingCtx, cancel := context.WithTimeout(ctx, timeout) defer cancel() From 37ef7578002562f34a0f9f2ce25c707ff80e585f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:11:07 -0400 Subject: [PATCH 022/108] fix(ingest): windowed reserve/publish/commit; 503 when dedupe is unavailable Each window of up to 256 records is prepared, reserved in one dedupe call, published in order and committed in one call. Deduped records are published under an idempotency key (Nats-Msg-Id), and each tenant's ingest stream keeps an explicit two-minute duplicate window, so a publish whose outcome is unknown leaves its claim to lapse and the retry's copy is dropped by the queue. A dedupe store that cannot answer is a 503 with Retry-After: 5. Fixes #384. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 1 + docs/src/content/docs/api.md | 12 +- docs/src/content/docs/architecture.md | 21 +- docs/src/content/docs/durability.md | 6 + docs/src/content/docs/sdk/reference.md | 2 +- docs/src/content/docs/settings-directory.mdx | 4 +- internal/api/ingest.go | 338 +++++++++----- internal/api/ingest_seams.go | 4 +- internal/api/ingest_test.go | 122 ++++-- internal/api/ingest_window_test.go | 435 +++++++++++++++++++ internal/dedupe/key.go | 9 + internal/dedupe/key_test.go | 33 ++ internal/mq/embedded.go | 17 +- internal/mq/embedded_test.go | 29 ++ internal/mq/mq.go | 14 + internal/testutil/mocks.go | 31 +- 17 files changed, 899 insertions(+), 183 deletions(-) create mode 100644 internal/api/ingest_window_test.go create mode 100644 internal/dedupe/key_test.go diff --git a/AGENTS.md b/AGENTS.md index 63615686..ffbd3c16 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` +- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal; `WithIdempotencyKey` makes a republish inside the queue's duplicate window a no-op), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) - **`policy/`** — Hasura-style access control, **role-first**: `TablePolicy` is `map[string]RolePermissions`, and a role's grant splits by operation into `SelectPermissions` (columns, row `filter`, aggregations, the `max_*` limits) and `InsertPermissions` (columns, `check`) — so a field only one side honors does not exist on the other. `Evaluate()` resolves ONE operation and leaves the other side **nil** (`Select *ResolvedSelect` / `Insert *ResolvedInsert`), which every accessor fails closed on — nil is "not resolved", distinct from an empty side, which is "unrestricted" (what the admin return builds). Claim templating (`{{ jwt.claim.path }}`) resolves during that call. Policies come from `Source`, a `func() *Policy` read per call (`settings.Store.Policy` in production, `Static(p)` in tests) @@ -58,7 +58,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. 6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format`, or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. -8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase — `Reserve` → publish → `Commit`, or `Release` when the publish fails; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. +8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase, one call per phase per window of up to 256 records — `Reserve` → publish (under the id's idempotency key) → `Commit`, or `Release` when the publish definitely failed, while one whose outcome is unknown is left to lapse; a store that cannot answer is a `503` + `Retry-After`; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 56aa5a3a..041ef99e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,6 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and the retry after the lease is dropped by the queue if the first copy was stored. A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index dc79ed6b..11f9f1f0 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -272,8 +272,9 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now (not open, or a remote backend throttled or unreachable); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate. | +| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored (it remembers the key for two minutes), so the retry never stores a second copy. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | @@ -385,13 +386,14 @@ A `200` is returned whenever the body was read and the records were processed | 403 | `{"error":"forbidden"}` (empty-role variant: `forbidden: request has no role and no public default_role is configured`) | The resolved role lacks `insert` on the table (checked once, before any record) | | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | -| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest | -| 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. As for `publish failed`, the failing record's id is given back and the records before it keep theirs | -| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | +| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the records before it keep their ids, so a whole-batch retry reports those as duplicates; the failing record's id is left to lapse as on the single-object path, and the rest of its window's ids are given back | +| 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. The records before the refused one keep their ids, and its id and the rest of its window's are given back | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now; `Retry-After: 5`. Nothing in the window being reserved was published; the windows before it were, and keep their ids | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). Nothing in that record's window was published; the windows before it were | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] -A batch aborted partway (a `503`/`500`, a JSON-array syntax error, or an NDJSON line over the 10 MiB line bound, after some leading records were already published) re-publishes those leading records when the whole batch is retried. A whole-body read failure is **not** one of these: a `413`, or the `400 invalid request body` of an upload cut off in transit, is decided before any record is processed, so nothing is published — safe to retry, once split for a `413`. Enable deduplication if duplicate suppression matters — this is the same at-least-once property the single-object path already has (the SDK retries both on `503`). +A batch aborted partway (a `503`/`500`, a JSON-array syntax error, or an NDJSON line over the 10 MiB line bound, after some leading records were already published) re-publishes those leading records when the whole batch is retried. Records are published in windows of 256, in order: a read error or a dedupe failure drops the open window unpublished, so what an aborted batch published is the windows before it, plus, after a publish failure, the records of its window before the failing one. A whole-body read failure is **not** one of these: a `413`, or the `400 invalid request body` of an upload cut off in transit, is decided before any record is processed, so nothing is published — safe to retry, once split for a `413`. Enable deduplication if duplicate suppression matters — this is the same at-least-once property the single-object path already has (the SDK retries both on `503`). ::: **curl example (JSON array):** diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 74066932..eb660c17 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -80,7 +80,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. -- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup (the id reserved once the record is encoded, committed after the publish, released if the publish fails; an id another request holds answers `503` with the lease as `Retry-After`), and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). +- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates and encodes each record, and runs the records in windows of up to 256 (`ingestWindow`) through three phases: one dedupe `Reserve` for the window's ids, the publishes in record order (a deduped record under `mq.WithIdempotencyKey`, keyed by `dedupe.IdempotencyKey`), and one `Commit` of the published ids — a window is the unit of a dedupe round trip and of Pebble's commit `fsync`. An id another request holds answers `503` with the lease as `Retry-After`, a store that cannot answer (`dedupe.ErrUnavailable`) `503` with `Retry-After: 5`; a publish that fails at a record commits the ones before it and releases the rest, except that a failure other than `mq.ErrQueueFull` may have stored the event, so that record's claim is left to lapse and the idempotency key drops the retry's copy. Each row goes through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. @@ -147,10 +147,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. -- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. +- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape; `WithIdempotencyKey` marks a publish so that a second one carrying the same key inside the queue's duplicate window is dropped and reported as success. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`) and remembering idempotency keys for `EmbeddedDuplicateWindow` (two minutes, which a dedupe lease must not exceed), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline @@ -228,11 +228,16 @@ Client POST /v1/ingest?table={table} an unparseable value passes through verbatim for ClickHouse's parser to judge) → Optional dedupe: resolve the id (configurable ID field; a row missing it or setting it to null is published un-deduped + logged/counted, or rejected - under require_id); once the record is encoded, reserve (tenant, table, id): - a duplicate is skipped, an id another request holds → 503 + Retry-After - (the 30s lease) - → Publish to NATS JetStream (ingest.{tenant}.{table}) - → Commit the reserved id; on a failed publish, release it instead + under require_id) + → Encode the record; the steps below run per window of up to 256 records + → Reserve the window's (tenant, table, id) keys in one call: a duplicate is + skipped, an id another request holds → 503 + Retry-After (the 30s lease), + a store that cannot answer → 503 + Retry-After: 5 + → Publish each record to NATS JetStream (ingest.{tenant}.{table}), a deduped + one under its idempotency key + → Commit the published ids in one call; on a failed publish, commit the + records before it and release the rest (a publish whose outcome is unknown + keeps its claim until the lease lapses) → 200 OK returned immediately → (If the tenant's NATS stream is full, or not open: 503 + Retry-After header, the id released) diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 8e57d823..a0e8e36e 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -58,6 +58,12 @@ The strict guarantee translates well to managed cloud infrastructure — the pre The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) is that a single-threaded benchmark looks fine while a concurrent one is far worse — so always benchmark with multiple writers, and benchmark the guest **and** the host if virtualized. +## Deduplication: one more fsync per window + +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids are committed to the dedupe store, and on the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. + +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes. The retry that follows the lease therefore stores no second copy. That holds only while the lease is shorter than the stream's duplicate window. + ## Check your storage before you trust it Replicate JetStream's exact pattern — a 4 KiB write followed by a flush, in a tight loop — and report the percentiles. The numbers that matter are **p99** and **max**: those are your worst-case publish latency. diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index a7a44387..40ca0048 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -32,7 +32,7 @@ The SDK **never throws** for anything the server returns — all API errors come | 403 | `HTTP_403` | No | Insufficient permissions | | 404 | `HTTP_404` | No | Table, pipe, or tenant not found | | 500 | `HTTP_500` | Yes | Server error (retried per `maxRetries`) | -| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | +| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a dedupe store that cannot answer (`dedupe store unavailable`, `Retry-After: 5`), a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | | 0 | `NETWORK_ERROR` | Yes | Network failure (retried with exponential backoff) | | 0 | `ABORTED` | No | Request canceled via `AbortSignal` | | 0 | `SSE_CONNECT_ERROR` | No | Stream could not be started (e.g. a non-absolute `baseURL`) | diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 905ba816..a9bdeb82 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,8 +185,8 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index f4051516..acf8bec8 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -37,6 +37,11 @@ import ( // with the admin query handler — see internal/api/query.go. const maxReportedResults = 10000 +// ingestWindow is how many records a batch prepares before reserving, +// publishing and committing them together: one dedupe call per phase per +// window rather than per record, and at most one window of encoded rows held. +const ingestWindow = 256 + // IngestHandler handles POST /v1/ingest?table={table} type IngestHandler struct { // Registry yields the request tenant's schema registry. @@ -68,6 +73,8 @@ type IngestHandler struct { // tests can pin the cap-overflow path without allocating 16 MiB per run; not // a production tuning knob, hence unexported. Mirrors QueryHandler. maxRequestBytes int64 + // window overrides ingestWindow when > 0, for tests and benchmarks. + window int } func NewIngestHandler(registry RegistrySource, pub mq.Publisher) *IngestHandler { @@ -133,11 +140,14 @@ type recordReject struct { // requestAbort is a whole-request failure: this record and every one that // follows is refused. Both paths stop and return the status; the batch path -// abandons the remaining records rather than silently losing the tail. +// abandons the remaining records rather than silently losing the tail. What +// earlier windows published stays published, and with dedupe on stays +// committed, so a whole-batch retry reports those records as duplicates. // // Most causes are TRANSIENT system conditions, where abandoning the tail is what // makes the batch safe to retry: publish backpressure (503), a publish/marshal -// failure (500), a dedup backend error (500), an id another request holds (503). +// failure (500), a dedupe store that cannot answer (503) or fails (500), an id +// another request holds (503). // // One is not. An insert grant that resolved for the other operation is a 403 and // a caller/config bug — retrying cannot help. It aborts rather than rejecting @@ -322,16 +332,21 @@ func (h *IngestHandler) handleSingle( return } - dup, reject, abort := h.processRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) + rec, abort := h.prepareRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) + if abort == nil && rec.reject == nil { + window := []pendingRecord{rec} + abort = h.ingestWindow(ctx, store, table, scope, window) + rec = window[0] + } if abort != nil { writeAbort(w, abort) return } - if reject != nil { - writeJSONError(w, reject.Status, reject.Message) + if rec.reject != nil { + writeJSONError(w, rec.reject.Status, rec.reject.Message) return } - if dup { + if rec.duplicate { w.Header().Set("Content-Type", "application/json") _ = json.NewEncoder(w).Encode(map[string]bool{"duplicate": true}) return @@ -363,6 +378,24 @@ func (h *IngestHandler) handleBatch( checkGuard *recordReject, ) { result := batchResult{Results: []recordResult{}} + size := h.window + if size <= 0 { + size = ingestWindow + } + window := make([]pendingRecord, 0, min(size, 16)) + // flush runs the window's records through reserve → publish → commit and + // reports them in order; false when it aborted the request. + flush := func() bool { + if abort := h.ingestWindow(ctx, store, table, scope, window); abort != nil { + writeAbort(w, abort) + return false + } + for i := range window { + result.add(&window[i]) + } + window = window[:0] + return true + } for { data, err := rr.Next() @@ -372,8 +405,10 @@ func (h *IngestHandler) handleBatch( if err != nil { if rse, ok := errors.AsType[*recordSyntaxError](err); ok { result.Total++ - result.Failed++ - appendResult(&result, recordResult{Index: result.Total, Error: rse.Error()}) + window = append(window, pendingRecord{index: result.Total, reject: &recordReject{Message: rse.Error()}}) + if len(window) == size && !flush() { + return + } continue } // Unreachable while the body is buffered — a bytes.Reader cannot produce @@ -392,26 +427,22 @@ func (h *IngestHandler) handleBatch( } result.Total++ - idx := result.Total - dup, reject, abort := h.processRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) + rec, abort := h.prepareRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) if abort != nil { // Whole-request failure: surface the status rather than recording a // request-scoped condition as per-record loss (see requestAbort). + // Nothing in the open window has been published. writeAbort(w, abort) return } - if reject != nil { - result.Failed++ - appendResult(&result, recordResult{Index: idx, Error: reject.Message}) - continue - } - if dup { - result.Duplicates++ - appendResult(&result, recordResult{Index: idx, Duplicate: true}) - continue + rec.index = result.Total + window = append(window, rec) + if len(window) == size && !flush() { + return } - result.Succeeded++ - appendResult(&result, recordResult{Index: idx, Ok: true}) + } + if len(window) > 0 && !flush() { + return } slog.InfoContext(ctx, "batch ingested", "table", table, @@ -421,12 +452,24 @@ func (h *IngestHandler) handleBatch( _ = json.NewEncoder(w).Encode(result) } -// appendResult records a per-record outcome up to maxReportedResults. The -// batchResult counts are incremented by the caller and stay authoritative even -// when the Results slice is truncated. -func appendResult(result *batchResult, entry recordResult) { - if len(result.Results) < maxReportedResults { - result.Results = append(result.Results, entry) +// add counts rec's outcome and records it up to maxReportedResults; the counts +// stay authoritative when Results is truncated. Total is counted as records +// are read. +func (r *batchResult) add(rec *pendingRecord) { + entry := recordResult{Index: rec.index} + switch { + case rec.reject != nil: + r.Failed++ + entry.Error = rec.reject.Message + case rec.duplicate: + r.Duplicates++ + entry.Duplicate = true + default: + r.Succeeded++ + entry.Ok = true + } + if len(r.Results) < maxReportedResults { + r.Results = append(r.Results, entry) } } @@ -523,19 +566,30 @@ func (h *IngestHandler) policyCheckGuard( } } -// processRecord runs the per-record pipeline shared by the single-object and -// batch ingest paths: schema validation → column/check permission enforcement -// (with claim-derived auto-injection) → optional dedup → publish. The -// table-level insert grant is checked once by the caller before any record is -// processed, so perms here drives only the per-column and per-row checks (it is -// nil when no policy store is configured). data may be mutated to auto-inject -// check-clause values. +// pendingRecord is one record between prepareRecord and its outcome. +type pendingRecord struct { + index int // 1-based position in a batch + reject *recordReject // non-nil: the record is bad and is not published + payload []byte // the encoded envelope to publish + // key is the record's dedupe identity, nil when it is published + // un-deduped; claim is Reserve's answer for it. + key *dedupe.Key + claim dedupe.Claim + duplicate bool +} + +// prepareRecord runs the per-record half of the pipeline shared by the +// single-object and batch ingest paths: schema validation → column/check +// permission enforcement (with claim-derived auto-injection) → timestamp +// canonicalization → dedupe id resolution → encoding. Reserving, publishing +// and committing happen per window, in ingestWindow. The table-level insert +// grant is checked once by the caller before any record is processed, so perms +// here drives only the per-column and per-row checks (it is nil when no policy +// store is configured). data may be mutated to auto-inject check-clause values. // -// Exactly one of the outcomes is meaningful per call: -// - duplicate true: the record was skipped by dedup (reject/abort nil). -// - reject non-nil: the record is bad; the rest of a batch may still proceed. -// - abort non-nil: a whole-request failure; the caller stops and returns it. -func (h *IngestHandler) processRecord( +// A record the rest of a batch may proceed past comes back with reject set; +// abort non-nil is a whole-request failure the caller stops and returns. +func (h *IngestHandler) prepareRecord( ctx context.Context, store *settings.Store, table, scope string, @@ -545,10 +599,10 @@ func (h *IngestHandler) processRecord( data map[string]any, now time.Time, checkGuard *recordReject, -) (duplicate bool, reject *recordReject, abort *requestAbort) { +) (rec pendingRecord, abort *requestAbort) { if err := h.validator().Validate(schema, data); err != nil { slog.WarnContext(ctx, "schema validation failed", "error", err, "table", table) - return false, &recordReject{Status: http.StatusBadRequest, Message: err.Error()}, nil + return pendingRecord{reject: &recordReject{Status: http.StatusBadRequest, Message: err.Error()}}, nil } // DEEP AUTH: column-level allow/deny + check clauses. @@ -556,10 +610,10 @@ func (h *IngestHandler) processRecord( for col := range data { if !perms.IsColumnAllowed(col, true) { slog.WarnContext(ctx, "column insertion forbidden", "column", col, "role", role) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusForbidden, Message: fmt.Sprintf("column %q not allowed for insert", col), - }, nil + }}, nil } } // Through the accessor, not a bare read. The check loop iterates a side's @@ -580,7 +634,7 @@ func (h *IngestHandler) processRecord( // permission failures for one mis-wired grant. slog.ErrorContext(ctx, "insert checks consulted on a grant resolved for another operation", "table", table, "role", role) - return false, nil, &requestAbort{ + return pendingRecord{}, &requestAbort{ Status: http.StatusForbidden, Message: "insert permissions were not resolved for this request", } @@ -595,7 +649,7 @@ func (h *IngestHandler) processRecord( // a record that supplies the column fails schema validation first with // a different message, and a batch should report each its own cause. if checkGuard != nil { - return false, checkGuard, nil + return pendingRecord{reject: checkGuard}, nil } // A []any value is an _in check: the inserted value must be present and // one of the allowed set. Unlike the scalar _eq case there is no single @@ -604,10 +658,10 @@ func (h *IngestHandler) processRecord( actual, ok := data[col] if !ok || !h.checker().InSet(actual, set) { slog.WarnContext(ctx, "check clause failed", "column", col, "allowed", set, "actual", actual, "present", ok) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusForbidden, Message: fmt.Sprintf("check failed for column %q", col), - }, nil + }}, nil } continue } @@ -624,10 +678,10 @@ func (h *IngestHandler) processRecord( // reading the token's own JSON type didn't give it. if !h.checker().Matches(actual, requiredVal) { slog.WarnContext(ctx, "check clause failed", "column", col, "expected", requiredVal, "actual", actual) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusForbidden, Message: fmt.Sprintf("check failed for column %q", col), - }, nil + }}, nil } } else { // Auto-inject the required value if not provided — as a plain @@ -652,23 +706,22 @@ func (h *IngestHandler) processRecord( // always states them, so no compiled fallback is needed), so a reload // lands at a record boundary. A Deduplicator without a settings source is // a wiring bug, not a mode — main wires both or neither. The id is claimed - // only once the record is encoded, so nothing but the publish can fail - // while the claim is held. - var dedupKey *dedupe.Key + // in ingestWindow, once every record of the window is encoded, so nothing + // but the publish can fail while the claim is held. if h.Dedup != nil && h.DedupeSettings != nil { if enabled, idField, requireID := h.DedupeSettings(store, table); enabled { // An explicit null is as missing as an absent key (#370): fmt.Sprint // would make every null "", one id for every such record. if idVal, ok := data[idField]; ok && idVal != nil { - dedupKey = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} + rec.key = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} } else { dedupeMissingIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) if requireID { slog.WarnContext(ctx, "dedupe id_field missing or null; rejecting", "id_field", idField, "table", table) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusBadRequest, Message: fmt.Sprintf("missing dedupe id field %q", idField), - }, nil + }}, nil } slog.WarnContext(ctx, "dedupe id_field missing or null; publishing without idempotency", "id_field", idField, "table", table) } @@ -683,7 +736,7 @@ func (h *IngestHandler) processRecord( row, err := ingest.EncodeCompactRow(cols, data) if err != nil { slog.ErrorContext(ctx, "failed to encode compact row", "error", err, "table", table) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} + return pendingRecord{}, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} } evt := ingest.EventMessage{ @@ -695,104 +748,173 @@ func (h *IngestHandler) processRecord( Row: row, } - payload, err := json.Marshal(evt) + rec.payload, err = json.Marshal(evt) if err != nil { slog.ErrorContext(ctx, "failed to marshal event message", "error", err) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} + return pendingRecord{}, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} } + return rec, nil +} +// ingestWindow reserves, publishes and commits one window of prepared +// records, in three phases: one Reserve for every keyed record, the publishes +// in record order, one Commit for every claim published. Rejected and +// duplicate records are skipped. It sets each record's outcome and returns an +// abort when the request must stop; what the window published before a failure +// is committed first, so the retry reports it as duplicates (see publishFailed). +func (h *IngestHandler) ingestWindow(ctx context.Context, store *settings.Store, table, scope string, recs []pendingRecord) *requestAbort { var dd dedupe.Deduplicator - var claims []dedupe.Claim - if dedupKey != nil { + var keyed []int + for i := range recs { + if recs[i].reject == nil && recs[i].key != nil { + keyed = append(keyed, i) + } + } + if len(keyed) > 0 { dd = h.Dedup(store) - var duplicate bool - var abort *requestAbort - claims, duplicate, abort = h.reserve(ctx, dd, *dedupKey) - if duplicate || abort != nil { - return duplicate, nil, abort + if abort := h.reserve(ctx, dd, table, recs, keyed); abort != nil { + return abort } } - slog.DebugContext(ctx, "publishing event to the ingest queue", "table", table, "scope", scope) - if err := h.Publisher.Publish(ctx, mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope}, payload); err != nil { - // The record is not in the queue, so its id goes back: the client's - // retry must not read as a duplicate of it (#384). - releaseClaims(ctx, dd, claims) - if errors.Is(err, mq.ErrQueueFull) { - slog.WarnContext(ctx, "ingest queue is full", "tenant", store.Tenant(), "error", err, "table", table, "scope", scope) - return false, nil, &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} + topic := mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope} + for i := range recs { + rec := &recs[i] + if rec.reject != nil || rec.duplicate { + continue + } + var opts []mq.PublishOpt + if rec.claim.Status == dedupe.Claimed { + // The retry of an uncertain publish carries the same id, so the + // queue drops its copy if the first one landed. + opts = append(opts, mq.WithIdempotencyKey(dedupe.IdempotencyKey(store.Tenant(), rec.claim.Key))) + } + if err := h.Publisher.Publish(ctx, topic, rec.payload, opts...); err != nil { + return h.publishFailed(ctx, dd, topic, recs, i, err) } - slog.ErrorContext(ctx, "failed to publish to the ingest queue", "tenant", store.Tenant(), "error", err, "table", table, "scope", scope) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "publish failed"} } - commitClaims(ctx, dd, claims, table) - return false, nil, nil + commitClaims(ctx, dd, claimedIn(recs), table) + return nil } -// reserve claims key for one record. A duplicate skips the record; a key -// another request holds aborts with 503 and the lease as Retry-After, since -// that request's outcome decides this one's. ErrDisabled — a reload switched -// the store off after the settings snapshot was read — publishes un-deduped, -// as a record under the other setting would have been. -func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, key dedupe.Key) (claims []dedupe.Claim, duplicate bool, abort *requestAbort) { +// reserve claims the keys of recs[keyed] in one call and records each answer. +// A duplicate is skipped. A key another request holds releases the window's +// claims and aborts with 503 and the lease as Retry-After, since that +// request's outcome decides this one's. A store that cannot answer now is a +// 503 too; nothing in the window has been published. ErrDisabled — a reload +// switched the store off after the settings snapshot was read — publishes the +// window un-deduped, as records under the other setting would have been. +func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, table string, recs []pendingRecord, keyed []int) *requestAbort { lease := h.DedupeLease if lease <= 0 { lease = dedupe.DefaultLease } - claims, err := dd.Reserve(ctx, []dedupe.Key{key}, lease) + keys := make([]dedupe.Key, len(keyed)) + for j, i := range keyed { + keys[j] = *recs[i].key + } + claims, err := dd.Reserve(ctx, keys, lease) switch { case errors.Is(err, dedupe.ErrDisabled): // The counter carries the signal (a burst is a reload; a steady rate // is the store and settings out of step), so the line is Debug rather // than a WARN per record. - dedupeDisabledCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", key.Table))) - slog.DebugContext(ctx, "dedupe switched off mid-reload; publishing without idempotency", "event_id", key.ID, "table", key.Table) - return nil, false, nil + dedupeDisabledCounter.Add(ctx, int64(len(keys)), metric.WithAttributes(attribute.String("table", table))) + slog.DebugContext(ctx, "dedupe switched off mid-reload; publishing without idempotency", "records", len(keys), "table", table) + return nil + case errors.Is(err, dedupe.ErrUnavailable): + slog.WarnContext(ctx, "dedupe store unavailable", "error", err, "table", table) + return &requestAbort{Status: http.StatusServiceUnavailable, Message: "dedupe store unavailable", RetryAfter: "5"} case err != nil: - slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "event_id", key.ID, "table", key.Table) - return nil, false, &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} - } - switch claims[0].Status { - case dedupe.Duplicate: - slog.InfoContext(ctx, "duplicate event skipped", "event_id", key.ID, "table", key.Table) - return nil, true, nil - case dedupe.InFlight: - slog.InfoContext(ctx, "event id in flight in another request", "event_id", key.ID, "table", key.Table) - return nil, false, &requestAbort{ + slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "table", table) + return &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} + } + var held *dedupe.Key + for j, i := range keyed { + recs[i].claim = claims[j] + switch claims[j].Status { + case dedupe.Duplicate: + recs[i].duplicate = true + slog.InfoContext(ctx, "duplicate event skipped", "event_id", keys[j].ID, "table", table) + case dedupe.InFlight: + if held == nil { + held = &keys[j] + } + case dedupe.Claimed: + } + } + if held != nil { + releaseClaims(ctx, dd, claimedIn(recs)) + slog.InfoContext(ctx, "event id in flight in another request", "event_id", held.ID, "table", table) + return &requestAbort{ Status: http.StatusServiceUnavailable, Message: "a request with the same dedupe id is in flight", RetryAfter: strconv.Itoa(int(math.Ceil(lease.Seconds()))), } - case dedupe.Claimed: } - return claims, false, nil + return nil +} + +// publishFailed settles a window whose publish failed at recs[k] and returns +// the abort. The records before k are queued, so their ids are committed. A +// definite failure — ErrQueueFull, the broker refused the event — releases k's +// id and the rest, so the client's retry publishes them (#384). Any other +// failure may have stored the event before failing, so k's claim is left to +// lapse with its lease instead: a retry before then answers in-flight, and one +// after republishes under the same idempotency key, which the queue drops if +// the first copy landed. The records after k were never sent and are released. +func (h *IngestHandler) publishFailed(ctx context.Context, dd dedupe.Deduplicator, topic mq.Topic, recs []pendingRecord, k int, err error) *requestAbort { + definite := errors.Is(err, mq.ErrQueueFull) + commitClaims(ctx, dd, claimedIn(recs[:k]), topic.Table) + after := k + 1 + if definite { + after = k + } + releaseClaims(ctx, dd, claimedIn(recs[after:])) + if definite { + slog.WarnContext(ctx, "ingest queue is full", "tenant", topic.Tenant, "error", err, "table", topic.Table, "scope", topic.Scope) + return &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} + } + slog.ErrorContext(ctx, "failed to publish to the ingest queue", "tenant", topic.Tenant, "error", err, "table", topic.Table, "scope", topic.Scope) + return &requestAbort{Status: http.StatusInternalServerError, Message: "publish failed"} +} + +// claimedIn is the Claimed claims among recs. +func claimedIn(recs []pendingRecord) []dedupe.Claim { + var out []dedupe.Claim + for i := range recs { + if recs[i].claim.Status == dedupe.Claimed { + out = append(out, recs[i].claim) + } + } + return out } -// commitClaims makes a published record's id a duplicate. A failure does not -// fail the record — it is in the queue — so it is logged and counted, and -// the claim lapses after its lease. +// commitClaims makes published records' ids duplicates. A failure does not +// fail the records — they are in the queue — so it is logged and counted, and +// the claims lapse after their lease. func commitClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim, table string) { if len(claims) == 0 { return } - // The record is queued whatever the request's context does next. + // The records are queued whatever the request's context does next. err := dd.Commit(context.WithoutCancel(ctx), claims, 0) switch { case err == nil, errors.Is(err, dedupe.ErrDisabled): default: - dedupeCommitFailedCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) - slog.ErrorContext(ctx, "dedupe commit failed after publish; the id lapses with its lease", "error", err, "table", table) + dedupeCommitFailedCounter.Add(ctx, int64(len(claims)), metric.WithAttributes(attribute.String("table", table))) + slog.ErrorContext(ctx, "dedupe commit failed after publish; the ids lapse with their lease", "error", err, "table", table, "records", len(claims)) } } -// releaseClaims gives claims back after a failed publish. A failure is only -// logged: the claim lapses with its lease either way. +// releaseClaims gives back claims whose records were not published. A failure +// is only logged: the claims lapse with their lease either way. func releaseClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim) { if len(claims) == 0 { return } if err := dd.Release(context.WithoutCancel(ctx), claims); err != nil && !errors.Is(err, dedupe.ErrDisabled) { - slog.WarnContext(ctx, "dedupe release failed; the id lapses with its lease", "error", err) + slog.WarnContext(ctx, "dedupe release failed; the ids lapse with their lease", "error", err) } } diff --git a/internal/api/ingest_seams.go b/internal/api/ingest_seams.go index d95fdb58..210f2ba5 100644 --- a/internal/api/ingest_seams.go +++ b/internal/api/ingest_seams.go @@ -20,7 +20,7 @@ import ( // return would invite a caller to change that. // // The two are one interface because they are one contract — "what this schema -// says about this record" — evaluated at two points in processRecord that must +// says about this record" — evaluated at two points in prepareRecord that must // stay apart: the insert-check block sits between them deliberately, so checks // keep pre-#372 semantics. type RecordValidator interface { @@ -55,7 +55,7 @@ func (h *IngestHandler) validator() RecordValidator { // InsertChecker decides whether a record's value satisfies a policy check // clause. Matches answers the scalar `_eq` form (the required value), InSet the // `_in` form (set membership). It never sees a record as a whole: the -// auto-injection of a missing check value stays in processRecord, where the +// auto-injection of a missing check value stays in prepareRecord, where the // ordering against validation and canonicalization is load-bearing. type InsertChecker interface { Matches(actual, required any) bool diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index f08d2eb0..57993333 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -660,7 +660,7 @@ func TestIngest_Policy_CheckIn_Absent_FailsClosed(t *testing.T) { // TestIngest_Policy_CheckIn_AbsentClaim_FailsClosed locks the typed-nil []any // path behind an _in check: when the claim itself is absent, resolveInValues -// returns a typed-nil []any, which must still assert as []any in processRecord +// returns a typed-nil []any, which must still assert as []any in prepareRecord // (entering the membership branch) so the column is rejected — never treated as a // scalar _eq value and auto-injected. The sibling _Absent test omits the column // with the claim present; this one drops the claim too. Guards #224 fail-closed. @@ -1825,8 +1825,8 @@ func TestIngest_JSONArray_SyntaxError_Fatal(t *testing.T) { h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) // A structural syntax error desyncs the decoder — the whole request fails - // (400), unlike a per-element type error. The leading good element may have - // already published (at-least-once on retry). + // (400), unlike a per-element type error. The leading good element is still + // in the open window, which is dropped unpublished. req := rawIngestRequest(t, "clicks", "application/json", `[{"page":"/a"}, {bad]`) w := httptest.NewRecorder() h.Handle(w, withTenant(req)) @@ -1834,7 +1834,7 @@ func TestIngest_JSONArray_SyntaxError_Fatal(t *testing.T) { assert.Equal(t, http.StatusBadRequest, w.Code) assert.Contains(t, w.Body.String(), "invalid json") testutil.AssertJSONErrorResponse(t, w) - assert.Len(t, pub.Messages, 1) // the leading record published before the abort + assert.Empty(t, pub.Messages) } func TestIngest_JSONArray_Truncated_Fatal(t *testing.T) { @@ -2347,7 +2347,7 @@ func TestIngest_Dedup_DisabledMidReload(t *testing.T) { // discovery.Validate accepts `{}` here because every column is nullable or // defaulted. I previously asserted this path was unreachable, having tested only // against a schema with a required column; it is not. -func TestProcessRecord_UnresolvedInsertSideAborts(t *testing.T) { +func TestPrepareRecord_UnresolvedInsertSideAborts(t *testing.T) { t.Parallel() schema := &discovery.TableSchema{ Name: "loose", @@ -2371,11 +2371,10 @@ func TestProcessRecord_UnresolvedInsertSideAborts(t *testing.T) { require.NoError(t, discovery.Validate(schema, map[string]any{}), "all-nullable/defaulted columns accept an empty record — this is what makes the read reachable") - dup, reject, abort := h.processRecord( + rec, abort := h.prepareRecord( context.Background(), testStore, "loose", "", schema, selectResolved, "viewer", map[string]any{}, time.Now(), nil) - assert.False(t, dup) - assert.Nil(t, reject, "a request-scoped condition must not be reported per record") + assert.Nil(t, rec.reject, "a request-scoped condition must not be reported per record") require.NotNil(t, abort, "an unresolved insert side must abort the request") assert.Equal(t, http.StatusForbidden, abort.Status) assert.Empty(t, abort.RetryAfter, "not a transient condition — retrying cannot help") @@ -2749,43 +2748,79 @@ func dedupHandler(t *testing.T, pub *testutil.MockPublisher, dedup dedupe.Dedupl return h } -// #384: a publish that fails gives the id back, so the retry the 503 asks for -// is published rather than skipped as a duplicate of a record that never -// reached the queue. +// #384: a publish the queue refused gives the id back, so the retry the 503 +// asks for is published rather than skipped as a duplicate of a record that +// never reached the queue. func TestIngest_Dedup_FailedPublishReleasesTheID(t *testing.T) { t.Parallel() - tests := []struct { - name string - err error - status int - }{ - {"backpressure", fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull), http.StatusServiceUnavailable}, - {"other failure", errors.New("connection reset"), http.StatusInternalServerError}, - } - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - pub := &testutil.MockPublisher{Err: tt.err} - dedup := testutil.NewMockDeduplicator() - h := dedupHandler(t, pub, dedup, false) - body := map[string]any{"page": "/home", "event_id": "e1"} + pub := &testutil.MockPublisher{Err: fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull)} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + body := map[string]any{"page": "/home", "event_id": "e1"} - w := httptest.NewRecorder() - h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) - require.Equal(t, tt.status, w.Code) - assert.False(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "released, not left to lapse") - - pub.Err = nil - w = httptest.NewRecorder() - h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) - require.Equal(t, http.StatusOK, w.Code) - assert.Contains(t, w.Body.String(), `"ok":true`, "the retry is published, not a duplicate") - assert.Len(t, pub.Published(), 1) - - w = httptest.NewRecorder() - h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) - assert.Contains(t, w.Body.String(), `"duplicate":true`, "and committed once published") - }) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "30", w.Header().Get("Retry-After")) + assert.False(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "released, not left to lapse") + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusOK, w.Code) + assert.Contains(t, w.Body.String(), `"ok":true`, "the retry is published, not a duplicate") + assert.Len(t, pub.Published(), 1) + + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + assert.Contains(t, w.Body.String(), `"duplicate":true`, "and committed once published") +} + +// A publish whose outcome is unknown may have stored the event, so its claim +// is neither released nor committed: it lapses with the lease, a retry before +// then answers in-flight, and the idempotency key covers one after. +func TestIngest_Dedup_UncertainPublishLeavesTheClaim(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: context.DeadlineExceeded} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + body := map[string]any{"page": "/home", "event_id": "e1"} + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusInternalServerError, w.Code) + assert.Contains(t, w.Body.String(), "publish failed") + assert.True(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "left to lapse") + assert.Empty(t, dedup.Released) + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "30", w.Header().Get("Retry-After")) + assert.Empty(t, pub.Published()) +} + +// A claimed record is published under its idempotency key; an un-deduped one +// carries none, so a producer's repeated ids are not dropped by the queue. +func TestIngest_Dedup_PublishCarriesTheIdempotencyKey(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, testutil.NewMockDeduplicator(), false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", + jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), + jsonLine(t, map[string]any{"page": "/b"}), + ))) + require.Equal(t, http.StatusOK, w.Code) + msgs := pub.Published() + require.Len(t, msgs, 2) + + want := mq.Headers{} + mq.WithIdempotencyKey(dedupe.IdempotencyKey(testStore.Tenant(), dedupe.Key{Table: "clicks", ID: "e1"}))(want) + for k, v := range want { + assert.Equal(t, v, msgs[0].Headers[k]) + assert.NotContains(t, msgs[1].Headers, k) } } @@ -2808,8 +2843,9 @@ func TestIngest_NDJSON_Dedup_PublishFailureMidBatch(t *testing.T) { w := httptest.NewRecorder() h.Handle(w, withTenant(batch())) require.Equal(t, http.StatusServiceUnavailable, w.Code) - require.Len(t, dedup.Released, 1) + require.Len(t, dedup.Released, 2, "the failing record and the rest of its window") assert.Equal(t, dedupe.Key{Table: "clicks", ID: "e2"}, dedup.Released[0].Key) + assert.Equal(t, dedupe.Key{Table: "clicks", ID: "e3"}, dedup.Released[1].Key) pub.Err = nil w = httptest.NewRecorder() diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go new file mode 100644 index 00000000..21f2e075 --- /dev/null +++ b/internal/api/ingest_window_test.go @@ -0,0 +1,435 @@ +package api + +import ( + "context" + "errors" + "fmt" + "net/http" + "net/http/httptest" + "strings" + "sync" + "testing" + "time" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/testutil" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// eventLines is n NDJSON clicks records with ids e1..en. +func eventLines(t *testing.T, n int) []string { + t.Helper() + lines := make([]string, n) + for i := range n { + lines[i] = jsonLine(t, map[string]any{"page": "/p", "event_id": fmt.Sprintf("e%d", i+1)}) + } + return lines +} + +func clickKey(i int) dedupe.Key { return dedupe.Key{Table: "clicks", ID: fmt.Sprintf("e%d", i)} } + +// A batch is reserved, published and committed a window at a time: one +// Reserve and one Commit per window, whatever the batch size. +func TestIngest_Windows_OneReserveAndCommitPerWindow(t *testing.T) { + t.Parallel() + for _, n := range []int{1, 255, 256, 257, 600} { + t.Run(fmt.Sprint(n), func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, n)...))) + require.Equal(t, http.StatusOK, w.Code) + assert.Equal(t, n, decodeBatchResult(t, w).Succeeded) + windows := (n + ingestWindow - 1) / ingestWindow + assert.Equal(t, windows, dedup.Reserves) + assert.Equal(t, windows, dedup.Commits) + assert.Len(t, pub.Published(), n) + assert.True(t, dedup.Committed(clickKey(n))) + }) + } +} + +// A publish failing at record k settles its window: the records before k are +// committed, k is released when the queue refused it and left to lapse when +// the outcome is unknown, the rest of the window is released, and later +// windows are never reserved. A whole-batch retry after a refusal publishes +// every record exactly once. +func TestIngest_Windows_PublishFailureAtK(t *testing.T) { + t.Parallel() + const n = 600 + refused := fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull) + tests := []struct { + name string + k int + err error + status int + }{ + {"refused first record", 1, refused, http.StatusServiceUnavailable}, + {"refused mid first window", 100, refused, http.StatusServiceUnavailable}, + {"refused last of first window", 256, refused, http.StatusServiceUnavailable}, + {"refused first of second window", 257, refused, http.StatusServiceUnavailable}, + {"refused mid last window", 590, refused, http.StatusServiceUnavailable}, + {"uncertain mid first window", 100, context.DeadlineExceeded, http.StatusInternalServerError}, + {"uncertain mid second window", 400, context.DeadlineExceeded, http.StatusInternalServerError}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: tt.err, ErrAfter: tt.k - 1} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + lines := eventLines(t, n) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", lines...))) + require.Equal(t, tt.status, w.Code) + assert.Len(t, pub.Published(), tt.k-1) + + windowEnd := min((tt.k-1)/ingestWindow*ingestWindow+ingestWindow, n) + definite := errors.Is(tt.err, mq.ErrQueueFull) + var released []dedupe.Key + for _, c := range dedup.Released { + released = append(released, c.Key) + } + var wantReleased []dedupe.Key + for i := tt.k; i <= windowEnd; i++ { + if i > tt.k || definite { + wantReleased = append(wantReleased, clickKey(i)) + } + } + assert.Equal(t, wantReleased, released) + if tt.k > 1 { + assert.True(t, dedup.Committed(clickKey(1))) + assert.True(t, dedup.Committed(clickKey(tt.k-1)), "published before the failure") + } + assert.False(t, dedup.Committed(clickKey(tt.k))) + assert.Equal(t, !definite, dedup.Pending(clickKey(tt.k)), "an uncertain publish leaves its claim to lapse") + if windowEnd < n { + assert.False(t, dedup.Pending(clickKey(windowEnd+1)), "a later window is never reserved") + } + if !definite { + return + } + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", lines...))) + require.Equal(t, http.StatusOK, w.Code) + resp := decodeBatchResult(t, w) + assert.Equal(t, tt.k-1, resp.Duplicates) + assert.Equal(t, n-(tt.k-1), resp.Succeeded) + assert.Len(t, pub.Published(), n, "every record exactly once") + }) + } +} + +// A dedupe store that cannot answer now fails the request with 503 and a +// short Retry-After, which the SDK retries — not the 500 of a broken store. +// Earlier windows stay published and committed. +func TestIngest_Dedup_UnavailableIs503(t *testing.T) { + t.Parallel() + notOpen := dedupe.NewManaged(func() (dedupe.Deduplicator, error) { return nil, errors.New("disk gone") }) + require.Error(t, notOpen.Apply(true)) + throttled := testutil.NewMockDeduplicator() + throttled.Err = fmt.Errorf("%w: throttled", dedupe.ErrUnavailable) + secondWindow := testutil.NewMockDeduplicator() + secondWindow.Err, secondWindow.ErrAfter = throttled.Err, 1 + + tests := []struct { + name string + dedup dedupe.Deduplicator + n int + published int + }{ + {"store not open", notOpen, 1, 0}, + {"backend throttled", throttled, 3, 0}, + {"second window throttled", secondWindow, ingestWindow + 1, ingestWindow}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, tt.dedup, false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, tt.n)...))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "5", w.Header().Get("Retry-After")) + assert.Contains(t, w.Body.String(), "dedupe store unavailable") + assert.Len(t, pub.Published(), tt.published) + }) + } + t.Run("single object", func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, throttled, false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/", "event_id": "e1"}))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "5", w.Header().Get("Retry-After")) + assert.Empty(t, pub.Published()) + }) +} + +// One id held by another request stops its window before anything in it is +// published and gives back the window's other claims; windows before it stay +// committed. +func TestIngest_Windows_InFlightReleasesTheWindow(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + held := ingestWindow + 2 + dedup.Hold(clickKey(held)) + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, ingestWindow+3)...))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "30", w.Header().Get("Retry-After")) + assert.Len(t, pub.Published(), ingestWindow) + assert.True(t, dedup.Committed(clickKey(ingestWindow))) + for _, i := range []int{ingestWindow + 1, ingestWindow + 3} { + assert.False(t, dedup.Pending(clickKey(i)), "e%d released", i) + } + assert.True(t, dedup.Pending(clickKey(held)), "the other request's claim is untouched") +} + +// Rejects, duplicates and repeats keep their places in the results across +// windows, over both batch formats. +func TestIngest_Windows_OutcomesStayInOrder(t *testing.T) { + t.Parallel() + records := []map[string]any{ + {"page": "/a", "event_id": "e1"}, + {"page": "/b", "event_id": "e1"}, // repeat inside one window + {"page": "/c", "nope": 1}, // reject + {"page": "/d", "event_id": "e2"}, + {"page": "/e", "event_id": "e1"}, // repeat across windows + {"page": "/f"}, // no id: published un-deduped + } + requests := map[string]func() *http.Request{ + "ndjson": func() *http.Request { + lines := make([]string, len(records)) + for i, r := range records { + lines[i] = jsonLine(t, r) + } + return ndjsonRequest(t, "clicks", lines...) + }, + "json array": func() *http.Request { return ingestRequest(t, "clicks", records) }, + } + for name, req := range requests { + t.Run(name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + h.window = 3 + + w := httptest.NewRecorder() + h.Handle(w, withTenant(req())) + require.Equal(t, http.StatusOK, w.Code) + resp := decodeBatchResult(t, w) + assert.Equal(t, []recordResult{ + {Index: 1, Ok: true}, + {Index: 2, Duplicate: true}, + {Index: 3, Error: resp.Results[2].Error}, + {Index: 4, Ok: true}, + {Index: 5, Duplicate: true}, + {Index: 6, Ok: true}, + }, resp.Results) + assert.NotEmpty(t, resp.Results[2].Error) + assert.Equal(t, 6, resp.Total) + assert.Len(t, pub.Published(), 3) + assert.Equal(t, 2, dedup.Reserves) + }) + } +} + +// The embedded queue must remember an idempotency key for at least a lease: +// the retry of an uncertain publish lands after the lease, and only the queue's +// duplicate window drops its second copy. +func TestIngest_DedupeLeaseFitsTheDuplicateWindow(t *testing.T) { + t.Parallel() + assert.LessOrEqual(t, dedupe.DefaultLease, mq.EmbeddedDuplicateWindow) +} + +// faultyPublisher publishes through a real broker and fails the calls fail +// picks: before sending (the queue refused it) or after (the outcome unknown +// to the caller, though the event is stored). +type faultyPublisher struct { + mq.Publisher + mu sync.Mutex + calls int + fail func(call int) (sendFirst bool, err error) +} + +func (p *faultyPublisher) Publish(ctx context.Context, topic mq.Topic, data []byte, opts ...mq.PublishOpt) error { + p.mu.Lock() + p.calls++ + sendFirst, err := p.fail(p.calls) + p.mu.Unlock() + if err == nil || sendFirst { + if pubErr := p.Publisher.Publish(ctx, topic, data, opts...); pubErr != nil { + return pubErr + } + } + return err +} + +// realPipeline is an ingest handler over the embedded broker and Pebble +// store, with pub's faults in front of the broker, and a count of the events +// in the tenant's queue. +func realPipeline(t *testing.T, fail func(call int) (bool, error)) (*IngestHandler, func() int) { + t.Helper() + broker, err := mq.NewEmbedded(t.TempDir()) + require.NoError(t, err) + t.Cleanup(func() { _ = broker.Close() }) + require.NoError(t, broker.SetMaxBytes(t.Context(), testStore.Tenant(), 64<<20)) + store := dedupe.NewEmbedded(t.TempDir()).Tenant(testStore.Tenant()) + require.NoError(t, store.Apply(true)) + t.Cleanup(func() { _ = store.Close() }) + + h := dedupHandler(t, nil, store, false) + h.Publisher = &faultyPublisher{Publisher: broker, fail: fail} + count := func() int { + n := 0 + require.NoError(t, broker.ReplaySince(t.Context(), mq.Topic{Tenant: testStore.Tenant(), Table: "clicks"}, time.Time{}, + func([]byte) bool { n++; return true })) + return n + } + return h, count +} + +// #384 end to end: a publish the queue refused, then the client's retry, ends +// in exactly one event in the queue — and a later retry is a duplicate. +func TestIngest_Dedup_FailedPublishThenRetryIsOneEvent(t *testing.T) { + t.Parallel() + h, count := realPipeline(t, func(call int) (bool, error) { + if call == 1 { + return false, fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull) + } + return false, nil + }) + body := map[string]any{"page": "/home", "event_id": "e1"} + codes := make([]int, 3) + for i := range codes { + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + codes[i] = w.Code + if i == 2 { + assert.Contains(t, w.Body.String(), `"duplicate":true`) + } + } + assert.Equal(t, []int{http.StatusServiceUnavailable, http.StatusOK, http.StatusOK}, codes) + assert.Equal(t, 1, count()) +} + +// A publish that stored the event but reported a failure, then the client's +// retry: in-flight until the lease lapses, then republished under the same +// idempotency key, which the queue drops — one event, and the id committed. +func TestIngest_Dedup_UncertainPublishThenRetryIsOneEvent(t *testing.T) { + t.Parallel() + h, count := realPipeline(t, func(call int) (bool, error) { + if call == 1 { + return true, context.DeadlineExceeded + } + return false, nil + }) + h.DedupeLease = 300 * time.Millisecond + lines := eventLines(t, 3) + send := func() *httptest.ResponseRecorder { + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", lines...))) + return w + } + + w := send() + require.Equal(t, http.StatusInternalServerError, w.Code) + w = send() + require.Equal(t, http.StatusServiceUnavailable, w.Code, "the uncertain claim is still held") + assert.Equal(t, "1", w.Header().Get("Retry-After")) + + var last *httptest.ResponseRecorder + require.Eventually(t, func() bool { + last = send() + return last.Code == http.StatusOK + }, 5*time.Second, 50*time.Millisecond) + assert.Equal(t, 3, decodeBatchResult(t, last).Succeeded, "the lapsed claim is claimed again and republished") + assert.Equal(t, 3, count(), "the republished e1 was dropped by the queue") + + w = send() + require.Equal(t, http.StatusOK, w.Code) + assert.Equal(t, 3, decodeBatchResult(t, w).Duplicates) +} + +// countingDedup counts the Commits that reach a store: on Pebble each is one +// fsync. +type countingDedup struct { + dedupe.Deduplicator + mu sync.Mutex + commits int +} + +func (c *countingDedup) Commit(ctx context.Context, claims []dedupe.Claim, retention time.Duration) error { + c.mu.Lock() + c.commits++ + c.mu.Unlock() + return c.Deduplicator.Commit(ctx, claims, retention) +} + +// pebbleBatchHandler is a handler over a real Pebble store behind a Commit +// counter, publishing to a mock queue. +func pebbleBatchHandler(tb testing.TB, window int) (*IngestHandler, *countingDedup) { + tb.Helper() + store := dedupe.NewEmbedded(tb.TempDir()).Tenant(testStore.Tenant()) + require.NoError(tb, store.Apply(true)) + tb.Cleanup(func() { _ = store.Close() }) + counted := &countingDedup{Deduplicator: store} + h := NewIngestHandler(fixedRegistry(testRegistry(tb)), &testutil.MockPublisher{}) + h.Dedup = staticDedup(counted) + h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.window = window + return h, counted +} + +// Windows cut the per-record fsyncs on Pebble: a 1,000-record batch commits in +// four syncs rather than a thousand. +func TestIngest_Windows_OneSyncPerWindowOnPebble(t *testing.T) { + t.Parallel() + h, counted := pebbleBatchHandler(t, 0) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, 1000)...))) + require.Equal(t, http.StatusOK, w.Code) + assert.Equal(t, 4, counted.commits) +} + +// BenchmarkIngest_DedupBatchOnPebble compares a 1,000-record batch committed +// per record (window 1, the pre-window behavior) with the default window. +func BenchmarkIngest_DedupBatchOnPebble(b *testing.B) { + for _, window := range []int{1, ingestWindow} { + b.Run(fmt.Sprintf("window=%d", window), func(b *testing.B) { + h, counted := pebbleBatchHandler(b, window) + var body strings.Builder + iter := 0 + for b.Loop() { + iter++ + body.Reset() + for i := range 1000 { + fmt.Fprintf(&body, `{"page":"/p","event_id":"%d-%d"}`+"\n", iter, i) + } + req := httptest.NewRequestWithContext(context.Background(), http.MethodPost, "/v1/ingest?table=clicks", strings.NewReader(body.String())) + req.Header.Set("Content-Type", "application/x-ndjson") + w := httptest.NewRecorder() + h.Handle(w, withTenant(req)) + if w.Code != http.StatusOK { + b.Fatalf("status %d", w.Code) + } + } + b.ReportMetric(float64(counted.commits)/float64(iter), "syncs/op") + }) + } +} diff --git a/internal/dedupe/key.go b/internal/dedupe/key.go index 25ead7c2..892e7c3f 100644 --- a/internal/dedupe/key.go +++ b/internal/dedupe/key.go @@ -2,6 +2,7 @@ package dedupe import ( "crypto/sha256" + "encoding/hex" "errors" "fmt" "strings" @@ -69,3 +70,11 @@ func AppendKey(dst, prefix []byte, k Key) []byte { } return append(dst, k.ID...) } + +// IdempotencyKey is k's message id for the queue under tenant id: the first +// 128 bits of the stored key's SHA-256, in hex, so a republished record is +// recognised without its id riding in a header verbatim. +func IdempotencyKey(id tenant.ID, k Key) string { + sum := sha256.Sum256(AppendKey(nil, KeyPrefix(id), k)) + return hex.EncodeToString(sum[:16]) +} diff --git a/internal/dedupe/key_test.go b/internal/dedupe/key_test.go new file mode 100644 index 00000000..2f2981e2 --- /dev/null +++ b/internal/dedupe/key_test.go @@ -0,0 +1,33 @@ +package dedupe_test + +import ( + "strings" + "testing" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/stretchr/testify/assert" +) + +// The idempotency key is 32 hex characters, stable for one tenant, table and +// id, and different when any of the three differs — ids too long to store +// verbatim included. +func TestIdempotencyKey(t *testing.T) { + t.Parallel() + long := strings.Repeat("x", dedupe.MaxIDBytes+1) + base := dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: "e1"}) + assert.Regexp(t, `^[0-9a-f]{32}$`, base) + assert.Equal(t, base, dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: "e1"})) + + others := []string{ + dedupe.IdempotencyKey("globex", dedupe.Key{Table: "clicks", ID: "e1"}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "views", ID: "e1"}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: "e2"}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: long}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: long + "y"}), + } + seen := map[string]bool{base: true} + for _, k := range others { + assert.False(t, seen[k], "collision: %s", k) + seen[k] = true + } +} diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 830be900..7ccc6e3e 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -306,17 +306,24 @@ func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { } } +// EmbeddedDuplicateWindow is how long an ingest queue remembers a +// WithIdempotencyKey key. A dedupe lease must not exceed it: a claim left to +// lapse after an uncertain publish is republished once the lease ends, and +// only this window drops that second copy. +const EmbeddedDuplicateWindow = 2 * time.Minute + // ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard // append-only log; the Active Sweeper handles message purging. MaxBytes caps // the tenant's share of the disk. DiscardNew rejects new messages when full, // propagating backpressure to the upstream API — for this tenant alone. func ingestStreamConfig(id tenant.ID, maxBytes int64) jetstream.StreamConfig { return jetstream.StreamConfig{ - Name: ingestStreamName(id), - Subjects: []string{tenantSubjects(ingestPrefix, id)}, - Retention: jetstream.LimitsPolicy, - MaxBytes: maxBytes, - Discard: jetstream.DiscardNew, + Name: ingestStreamName(id), + Subjects: []string{tenantSubjects(ingestPrefix, id)}, + Retention: jetstream.LimitsPolicy, + MaxBytes: maxBytes, + Discard: jetstream.DiscardNew, + Duplicates: EmbeddedDuplicateWindow, } } diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 6e87ff7b..004938d0 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -141,6 +141,34 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { assert.Equal(t, []byte("x"), raw.Data) } +// A repeated idempotency key inside the duplicate window is dropped as a +// success, so an uncertain publish can be republished safely; a queue made +// before the window was set gets it on its next budget apply. +func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + old := ingestStreamConfig(tenant.Default, testBudget) + old.Duplicates = 0 + _, err := e.js.CreateStream(ctx, old) + require.NoError(t, err) + require.NoError(t, e.SetMaxBytes(ctx, tenant.Default, testBudget)) + require.Equal(t, EmbeddedDuplicateWindow, streamConfig(t, e, "INGEST_0").Duplicates) + + topic := Topic{Tenant: tenant.Default, Table: "t"} + require.NoError(t, e.Publish(ctx, topic, []byte("a"), WithIdempotencyKey("k1"))) + require.NoError(t, e.Publish(ctx, topic, []byte("a again"), WithIdempotencyKey("k1")), "a repeat is a success") + require.NoError(t, e.Publish(ctx, topic, []byte("b"), WithIdempotencyKey("k2"))) + require.NoError(t, e.Publish(ctx, topic, []byte("c"))) + + var got []string + require.NoError(t, e.ReplaySince(ctx, topic, time.Time{}, func(data []byte) bool { + got = append(got, string(data)) + return true + })) + assert.Equal(t, []string{"a", "b", "c"}, got) +} + // A tenant's first budget opens its queue: an ingest stream holding its // subjects alone at the budget, refusing when full, and a dead-letter stream // at a tenth of it, dropping its oldest when full. No other tenant gets one. @@ -155,6 +183,7 @@ func TestEmbeddedNATS_SetMaxBytes_OpensTheTenantsQueue(t *testing.T) { assert.Equal(t, []string{"ingest.acme.>"}, ingest.Subjects) assert.Equal(t, int64(testBudget), ingest.MaxBytes) assert.Equal(t, jetstream.DiscardNew, ingest.Discard) + assert.Equal(t, EmbeddedDuplicateWindow, ingest.Duplicates) dlq := streamConfig(t, e, "DLQ_acme") assert.Equal(t, []string{"dlq.acme.>"}, dlq.Subjects) assert.Equal(t, int64(testBudget)/10, dlq.MaxBytes) diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 3f1c45c1..97eac365 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -141,6 +141,20 @@ func WithHeader(key, value string) PublishOpt { } } +// idempotencyHeader carries WithIdempotencyKey's key: JetStream's own +// message-id header, which the stream deduplicates on. +const idempotencyHeader = "Nats-Msg-Id" + +// WithIdempotencyKey marks a publish with key: a second publish carrying the +// same key within the queue's duplicate window is dropped by the broker and +// reported as success, so republishing an event whose first publish had an +// unknown outcome stores it once. +func WithIdempotencyKey(key string) PublishOpt { + return func(h Headers) { + h.Set(idempotencyHeader, key) + } +} + // ErrQueueFull is returned by Publisher.Publish when the topic's tenant's // ingest queue refuses new events — it is at its byte budget, or the tenant // has no queue open yet — the backpressure signal the API turns into a 503 diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index cdb1e33e..624fe0ef 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -118,11 +118,15 @@ type MockDeduplicator struct { committed map[dedupe.Key]bool pending map[dedupe.Key]string tokens int - // Err, if set, fails Reserve; CommitErr and ReleaseErr fail their phase. + // Err, if set, fails Reserve — after ErrAfter calls have succeeded; + // CommitErr and ReleaseErr fail their phase. Err error + ErrAfter int CommitErr error ReleaseErr error Released []dedupe.Claim // every claim Release was given + // Calls to each phase, for tests that count round trips. + Reserves, Commits int } var _ dedupe.Deduplicator = (*MockDeduplicator)(nil) @@ -131,16 +135,21 @@ func NewMockDeduplicator() *MockDeduplicator { return &MockDeduplicator{committed: map[dedupe.Key]bool{}, pending: map[dedupe.Key]string{}} } +// Reserve answers Duplicate for a key repeated in one call, as Managed does. func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time.Duration) ([]dedupe.Claim, error) { - if m.Err != nil { - return nil, m.Err - } m.mu.Lock() defer m.mu.Unlock() + m.Reserves++ + if m.Err != nil && m.Reserves > m.ErrAfter { + return nil, m.Err + } claims := make([]dedupe.Claim, 0, len(keys)) + seen := make(map[dedupe.Key]bool, len(keys)) for _, k := range keys { + repeat := seen[k] + seen[k] = true switch { - case m.committed[k]: + case repeat, m.committed[k]: claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.Duplicate}) case m.pending[k] != "": claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.InFlight}) @@ -155,11 +164,12 @@ func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time. } func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ time.Duration) error { + m.mu.Lock() + defer m.mu.Unlock() + m.Commits++ if m.CommitErr != nil { return m.CommitErr } - m.mu.Lock() - defer m.mu.Unlock() for _, c := range claims { if c.Status == dedupe.Claimed { m.committed[c.Key] = true @@ -192,6 +202,13 @@ func (m *MockDeduplicator) Hold(k dedupe.Key) { m.pending[k] = "held" } +// Committed reports whether k was committed. +func (m *MockDeduplicator) Committed(k dedupe.Key) bool { + m.mu.Lock() + defer m.mu.Unlock() + return m.committed[k] +} + // Pending reports whether k is claimed and neither committed nor released. func (m *MockDeduplicator) Pending(k dedupe.Key) bool { m.mu.Lock() From 71ab82482ca057dac5a47b0ae2fd6efbd3927f39 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:48:41 -0400 Subject: [PATCH 023/108] fix(ingest): qualify the duplicate-window claims; steadier tests The idempotency key drops a retry only within two minutes of the first publish, and a 200 with a failed commit is counted, not committed: the docs now say so. The uncertain-publish test uses a 2s lease so a stall cannot lapse the claim early, and the mq test pins the explicit duplicate window rather than the server's matching default. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 4 ++-- docs/src/content/docs/settings-directory.mdx | 2 +- internal/api/ingest_window_test.go | 6 +++--- internal/mq/embedded_test.go | 6 ++++-- 7 files changed, 13 insertions(+), 11 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 041ef99e..f0dab22e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and the retry after the lease is dropped by the queue if the first copy was stored. A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 11f9f1f0..dd84c32e 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -274,7 +274,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now (not open, or a remote backend throttled or unreachable); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored (it remembers the key for two minutes), so the retry never stores a second copy. | +| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue remembers the key for two minutes after the first publish, so a retry inside that window stores no second copy (the SDK's, after the 30-second `Retry-After`, lands inside it); a later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index eb660c17..e50b269f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -80,7 +80,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. -- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates and encodes each record, and runs the records in windows of up to 256 (`ingestWindow`) through three phases: one dedupe `Reserve` for the window's ids, the publishes in record order (a deduped record under `mq.WithIdempotencyKey`, keyed by `dedupe.IdempotencyKey`), and one `Commit` of the published ids — a window is the unit of a dedupe round trip and of Pebble's commit `fsync`. An id another request holds answers `503` with the lease as `Retry-After`, a store that cannot answer (`dedupe.ErrUnavailable`) `503` with `Retry-After: 5`; a publish that fails at a record commits the ones before it and releases the rest, except that a failure other than `mq.ErrQueueFull` may have stored the event, so that record's claim is left to lapse and the idempotency key drops the retry's copy. Each row goes through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). +- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates and encodes each record, and runs the records in windows of up to 256 (`ingestWindow`) through three phases: one dedupe `Reserve` for the window's ids, the publishes in record order (a deduped record under `mq.WithIdempotencyKey`, keyed by `dedupe.IdempotencyKey`), and one `Commit` of the published ids — a window is the unit of a dedupe round trip and of Pebble's commit `fsync`. An id another request holds answers `503` with the lease as `Retry-After`, a store that cannot answer (`dedupe.ErrUnavailable`) `503` with `Retry-After: 5`; a publish that fails at a record commits the ones before it and releases the rest, except that a failure other than `mq.ErrQueueFull` may have stored the event, so that record's claim is left to lapse and the idempotency key drops the retry's copy if it comes within the stream's two-minute duplicate window. Each row goes through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index a0e8e36e..5f84c624 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -60,9 +60,9 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) ## Deduplication: one more fsync per window -With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids are committed to the dedupe store, and on the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes. The retry that follows the lease therefore stores no second copy. That holds only while the lease is shorter than the stream's duplicate window. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The window covers a prompt retry only while the lease is shorter than it. ## Check your storage before you trust it diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index a9bdeb82..d10be43e 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index 21f2e075..3d1dc931 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -339,7 +339,7 @@ func TestIngest_Dedup_UncertainPublishThenRetryIsOneEvent(t *testing.T) { } return false, nil }) - h.DedupeLease = 300 * time.Millisecond + h.DedupeLease = 2 * time.Second lines := eventLines(t, 3) send := func() *httptest.ResponseRecorder { w := httptest.NewRecorder() @@ -351,13 +351,13 @@ func TestIngest_Dedup_UncertainPublishThenRetryIsOneEvent(t *testing.T) { require.Equal(t, http.StatusInternalServerError, w.Code) w = send() require.Equal(t, http.StatusServiceUnavailable, w.Code, "the uncertain claim is still held") - assert.Equal(t, "1", w.Header().Get("Retry-After")) + assert.Equal(t, "2", w.Header().Get("Retry-After")) var last *httptest.ResponseRecorder require.Eventually(t, func() bool { last = send() return last.Code == http.StatusOK - }, 5*time.Second, 50*time.Millisecond) + }, 10*time.Second, 100*time.Millisecond) assert.Equal(t, 3, decodeBatchResult(t, last).Succeeded, "the lapsed claim is claimed again and republished") assert.Equal(t, 3, count(), "the republished e1 was dropped by the queue") diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 004938d0..fc7900b0 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -143,13 +143,15 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // A repeated idempotency key inside the duplicate window is dropped as a // success, so an uncertain publish can be republished safely; a queue made -// before the window was set gets it on its next budget apply. +// with another window gets this one on its next budget apply. func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { e := openEmbedded(t, t.TempDir()) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() + // Explicit rather than the server's default, which happens to match today. + require.Equal(t, EmbeddedDuplicateWindow, ingestStreamConfig(tenant.Default, testBudget).Duplicates) old := ingestStreamConfig(tenant.Default, testBudget) - old.Duplicates = 0 + old.Duplicates = 10 * time.Second _, err := e.js.CreateStream(ctx, old) require.NoError(t, err) require.NoError(t, e.SetMaxBytes(ctx, tenant.Default, testBudget)) From 6f7944aa8ac566c0dd15e2069af42324abe8381e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:53:16 -0400 Subject: [PATCH 024/108] fix(dedupe): attempt every Commit and Release chunk; keep the breaker honest Review round 1. Commit and Release no longer cancel their siblings on the first failure: the records are already published, and a claim left behind holds its id for a lease. A failed Reserve sends no put after the first failure and releases only puts it sent; a sibling cancelled by that failure no longer resets the breaker (and is counted as outcome "canceled"). Docs: TTL reclaims only lapsed claims until retention lands (#220), no future boot-key names, the breaker counts consecutive unavailable claims. The e2e coverage suite excludes dynamodb.go, which the e2e binary never runs; unit and integration cover it and the merged total counts it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 5 ++ AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 10 +-- internal/dedupe/dynamodb.go | 86 ++++++++++++--------- internal/dedupe/dynamodb_test.go | 103 ++++++++++++++++++++++++-- 7 files changed, 162 insertions(+), 48 deletions(-) diff --git a/.testcoverage.yml b/.testcoverage.yml index aff1a694..27052082 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,3 +73,8 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ + # The DynamoDB dedupe backend: the e2e binary runs Pebble dedupe, so + # this file measured 0% there and pulled e2e to 58.6%. The unit + # (fake API) and integration (dynamodb-local) suites cover it, and the + # merged total still counts it. + - ^internal/dedupe/dynamodb\.go$ diff --git a/AGENTS.md b/AGENTS.md index 84b0a5ee..29e5e72c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch); `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims, built and conformance-tested against dynamodb-local but not yet selectable at boot), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; built and conformance-tested against dynamodb-local but not yet selectable at boot), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index f31bb3c1..7e11d26d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3cee2210..c4d7fcf8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -125,7 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims within a second short-circuit `Reserve` for a second. `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 2f97ad13..9bea802e 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -424,10 +424,10 @@ The dedupe key now carries the table as well as the tenant ([#222](https://githu ## A shared dedupe table on DynamoDB :::note[Not selectable yet] -The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. The `dedupe.backend` boot key lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. +The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. A boot key to select it lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. ::: -Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend's `create_table` switch is refused unless an `endpoint` override is set, so it only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). +Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend refuses to create a table unless it is pointed at a custom endpoint, so table creation only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). What the backend requires of the table: @@ -438,7 +438,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet. Without TTL, though, expired items are never removed and storage keeps growing. The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: @@ -489,8 +489,8 @@ data "aws_iam_policy_document" "wavehouse_dedupe" { - **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. -- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. -- **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five failed claims within one second, the backend stops calling the table for a second and fails requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). +- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. +- **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five throttled or unreachable claims in a row within one second, the backend stops calling the table for a second and fails every tenant's dedupe requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). A duplicate or in-flight answer is not a failure and resets the count. - **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. ## Upgrading across the v2 ingest envelope diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 2ce30e6e..08fd0c66 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -260,7 +260,9 @@ func (d *Dynamo) call(ctx context.Context, op string, do func(context.Context) e start := time.Now() err := classify(op, do(ctx)) d.metrics.record(ctx, op, time.Since(start), err) - if op == opReserve { + // A request cancelled because a sibling failed says nothing about the + // table, and must not reset the breaker's count. + if op == opReserve && !errors.Is(err, context.Canceled) { d.breaker.record(err) } return err @@ -289,12 +291,19 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati exp := expiresAt(now, lease) claims := make([]Claim, len(keys)) tried := make([]Claim, len(keys)) + sent := make([]bool, len(keys)) + // The first failure cancels the puts not yet sent: the Reserve fails + // either way, and a throttled table should not take the rest. g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { token := newToken() tried[i] = Claim{Key: k, Status: Claimed, Token: token} g.Go(func() error { + if err := gctx.Err(); err != nil { + return err + } + sent[i] = true status, err := s.reserve(gctx, k, token, nowSec, exp) if err != nil { return err @@ -307,11 +316,11 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati }) } if err := g.Wait(); err != nil { - // A put that errored or never answered may still have landed; its - // token is known, and releasing a key it does not hold is a no-op. + // A put that was sent and errored may still have landed; its token + // is known, and releasing a key it does not hold is a no-op. var undo []Claim for i, c := range claims { - if c.Status == Claimed || c.Status == 0 { + if sent[i] && (c.Status == Claimed || c.Status == 0) { undo = append(undo, tried[i]) } } @@ -381,13 +390,12 @@ func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time } writes = append(writes, types.WriteRequest{PutRequest: &types.PutRequest{Item: item}}) } - g, gctx := errgroup.WithContext(ctx) - g.SetLimit(s.d.cfg.ReserveConcurrency) - for start := 0; start < len(writes); start += batchWriteMax { - chunk := writes[start:min(start+batchWriteMax, len(writes))] - g.Go(func() error { return s.commitChunk(gctx, chunk) }) - } - return g.Wait() + // Every chunk is attempted whatever another's fate: these records are + // already published, and an uncommitted id lets a retry publish again. + chunks := (len(writes) + batchWriteMax - 1) / batchWriteMax + return forEach(chunks, s.d.cfg.ReserveConcurrency, func(i int) error { + return s.commitChunk(ctx, writes[i*batchWriteMax:min((i+1)*batchWriteMax, len(writes))]) + }) } func (s *dynamoStore) commitChunk(ctx context.Context, writes []types.WriteRequest) error { @@ -428,30 +436,27 @@ func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { if len(claims) == 0 { return nil } - g, gctx := errgroup.WithContext(ctx) - g.SetLimit(s.d.cfg.ReserveConcurrency) - for _, c := range claims { - g.Go(func() error { - err := s.d.call(gctx, "delete_item", func(ctx context.Context) error { - _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ - TableName: &s.d.cfg.Table, - Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, - ConditionExpression: aws.String(condRelease), - ExpressionAttributeValues: map[string]types.AttributeValue{ - ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, - ":pending": &types.AttributeValueMemberN{Value: statePending}, - }, - }) - return err + // Every claim is attempted: one left behind holds its id for a lease. + return forEach(len(claims), s.d.cfg.ReserveConcurrency, func(i int) error { + c := claims[i] + err := s.d.call(ctx, "delete_item", func(ctx context.Context) error { + _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ + TableName: &s.d.cfg.Table, + Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, + ConditionExpression: aws.String(condRelease), + ExpressionAttributeValues: map[string]types.AttributeValue{ + ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, + ":pending": &types.AttributeValueMemberN{Value: statePending}, + }, }) - var gone *types.ConditionalCheckFailedException - if errors.As(err, &gone) { - return nil - } return err }) - } - return g.Wait() + var gone *types.ConditionalCheckFailedException + if errors.As(err, &gone) { + return nil + } + return err + }) } // Close is a no-op: the client is the Dynamo's, shared by every tenant. @@ -459,6 +464,19 @@ func (s *dynamoStore) Close() error { return nil } // expiresAt is t+d in epoch seconds rounded up, so a claim or commit never // ends before it was asked to: TTL attributes are whole seconds. +// forEach runs do for every index, at most limit at once, and joins the +// errors: one failure never stops the rest. +func forEach(n, limit int, do func(i int) error) error { + errs := make([]error, n) + var g errgroup.Group + g.SetLimit(limit) + for i := range n { + g.Go(func() error { errs[i] = do(i); return nil }) + } + _ = g.Wait() + return errors.Join(errs...) +} + func expiresAt(t time.Time, d time.Duration) int64 { end := t.Add(d) sec := end.Unix() @@ -575,7 +593,7 @@ type dynamoMetrics struct { func newDynamoMetrics() dynamoMetrics { meter := otel.Meter("wavehouse-dedupe") requests, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_requests_total", - metric.WithDescription("DynamoDB dedupe requests by operation and outcome (ok, condition_failed, unavailable, error)")) + metric.WithDescription("DynamoDB dedupe requests by operation and outcome (ok, condition_failed, unavailable, canceled, error)")) duration, _ := meter.Float64Histogram("wavehouse_dedupe_dynamodb_request_duration_seconds", metric.WithDescription("DynamoDB dedupe request latency, SDK retries included"), metric.WithUnit("s")) unprocessed, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_unprocessed_items_total", @@ -592,6 +610,8 @@ func (m dynamoMetrics) record(ctx context.Context, op string, took time.Duration case err == nil: case errors.As(err, &cond): outcome = "condition_failed" + case errors.Is(err, context.Canceled): + outcome = "canceled" case errors.Is(err, ErrUnavailable): outcome = "unavailable" default: diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 8648c938..e6276126 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -24,18 +24,18 @@ import ( // dynamodb-local (tests/integration); this is for the error paths it cannot // produce. type fakeDynamo struct { - put func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) + put func(context.Context, *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) batch func(*dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) del func(*dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) describe func() (*dynamodb.DescribeTableOutput, error) ttl func() (*dynamodb.DescribeTimeToLiveOutput, error) } -func (f *fakeDynamo) PutItem(_ context.Context, in *dynamodb.PutItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) { +func (f *fakeDynamo) PutItem(ctx context.Context, in *dynamodb.PutItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) { if f.put == nil { return &dynamodb.PutItemOutput{}, nil } - return f.put(in) + return f.put(ctx, in) } func (f *fakeDynamo) BatchWriteItem(_ context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { @@ -74,7 +74,12 @@ func apiErr(code string, fault smithy.ErrorFault) error { func openFake(t *testing.T, f *fakeDynamo) (*Dynamo, Deduplicator) { t.Helper() - d := newDynamo(f, DynamoConfig{Table: "dedupe"}) + return openFakeWith(t, f, DynamoConfig{Table: "dedupe"}) +} + +func openFakeWith(t *testing.T, f *fakeDynamo, cfg DynamoConfig) (*Dynamo, Deduplicator) { + t.Helper() + d := newDynamo(f, cfg) d.commitBackoff = func(int) time.Duration { return 0 } m := d.Tenant("acme") require.NoError(t, m.Apply(true)) @@ -125,7 +130,7 @@ func TestClassify(t *testing.T) { func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { t.Parallel() - _, m := openFake(t, &fakeDynamo{put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + _, m := openFake(t, &fakeDynamo{put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) switch id[len(id)-1] { case 'd': @@ -150,7 +155,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { putTokens := map[string]string{} var released []string _, m := openFake(t, &fakeDynamo{ - put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) mu.Lock() putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) @@ -252,7 +257,7 @@ func TestDynamo_BreakerShortCircuitsReserve(t *testing.T) { var puts atomic.Int64 var down atomic.Bool down.Store(true) - d, m := openFake(t, &fakeDynamo{put: func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + d, m := openFake(t, &fakeDynamo{put: func(context.Context, *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { puts.Add(1) if down.Load() { return nil, &types.ProvisionedThroughputExceededException{} @@ -355,3 +360,87 @@ func TestExpiresAt(t *testing.T) { assert.Equal(t, int64(102), expiresAt(base, 1500*time.Millisecond), "rounded up: never ends early") assert.Equal(t, int64(102), expiresAt(base.Add(time.Nanosecond), time.Second)) } + +func idOf(av types.AttributeValue) string { + b := av.(*types.AttributeValueMemberB).Value + return string(b[len(b)-2:]) +} + +func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { + t.Parallel() + var mu sync.Mutex + written := 0 + _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + reqs := in.RequestItems["dedupe"] + if idOf(reqs[0].PutRequest.Item[attrKey]) == "00" { + return nil, &types.InternalServerError{} + } + mu.Lock() + written += len(reqs) + mu.Unlock() + return &dynamodb.BatchWriteItemOutput{}, nil + }}) + var claims []Claim + for i := range 3 * batchWriteMax { + claims = append(claims, Claim{Key: keys(fmt.Sprintf("%02d", i))[0], Status: Claimed, Token: "t"}) + } + require.ErrorIs(t, m.Commit(t.Context(), claims, 0), ErrUnavailable) + assert.Equal(t, 2*batchWriteMax, written, "a failed chunk does not cancel the others: their records are published") +} + +func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { + t.Parallel() + var deletes atomic.Int64 + _, m := openFakeWith(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + deletes.Add(1) + if idOf(in.Key[attrKey]) == "k0" { + return nil, &types.ProvisionedThroughputExceededException{} + } + return &dynamodb.DeleteItemOutput{}, nil + }}, DynamoConfig{Table: "dedupe", ReserveConcurrency: 1}) + var claims []Claim + for _, k := range keys("k0", "k1", "k2", "k3") { + claims = append(claims, Claim{Key: k, Status: Claimed, Token: "t"}) + } + require.ErrorIs(t, m.Release(t.Context(), claims), ErrUnavailable) + assert.Equal(t, int64(4), deletes.Load()) +} + +// One throttled put in a multi-key Reserve: the unsent puts are neither sent +// nor released, and the cancelled siblings do not reset the breaker. +func TestDynamo_FailedMultiKeyReserve(t *testing.T) { + t.Parallel() + var mu sync.Mutex + var put, released []string + _, m := openFakeWith(t, &fakeDynamo{ + put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := idOf(in.Item[attrKey]) + mu.Lock() + put = append(put, id) + mu.Unlock() + if id == "k0" { + return nil, &types.ProvisionedThroughputExceededException{} + } + <-ctx.Done() + return nil, ctx.Err() + }, + del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + mu.Lock() + released = append(released, idOf(in.Key[attrKey])) + mu.Unlock() + return nil, &types.ConditionalCheckFailedException{} + }, + }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) + ks := keys("k0", "k1", "k2", "k3", "k4", "k5", "k6", "k7") + for range breakerTrips { + _, err := m.Reserve(t.Context(), ks, time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + require.NotErrorIs(t, err, errBreakerOpen) + } + _, err := m.Reserve(t.Context(), ks, time.Minute) + require.ErrorIs(t, err, errBreakerOpen, "the cancelled siblings did not reset the count") + mu.Lock() + defer mu.Unlock() + assert.ElementsMatch(t, put, released, "exactly the sent puts are released") + assert.Less(t, len(put), breakerTrips*len(ks), "unsent puts were never sent") +} From ff047d2d374a1cadce7357c53571e2cae24c4cb9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:16:31 -0400 Subject: [PATCH 025/108] test(dedupe): make the Dynamo failure tests fail without their fix Review round 2. The fakes' BatchWriteItem and DeleteItem honour their context, and the Commit test runs one chunk at a time, so reverting Commit/Release to cancel-on-first-error fails both tests (checked against 108499f4). The failed-Reserve test asserts the sent puts are released rather than every key, which the unsent-put cutoff made flaky. forEach no longer sits between expiresAt and its doc comment. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dynamodb.go | 4 +-- internal/dedupe/dynamodb_test.go | 43 +++++++++++++++++++++----------- 2 files changed, 30 insertions(+), 17 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 08fd0c66..eb1c8383 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -462,8 +462,6 @@ func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { // Close is a no-op: the client is the Dynamo's, shared by every tenant. func (s *dynamoStore) Close() error { return nil } -// expiresAt is t+d in epoch seconds rounded up, so a claim or commit never -// ends before it was asked to: TTL attributes are whole seconds. // forEach runs do for every index, at most limit at once, and joins the // errors: one failure never stops the rest. func forEach(n, limit int, do func(i int) error) error { @@ -477,6 +475,8 @@ func forEach(n, limit int, do func(i int) error) error { return errors.Join(errs...) } +// expiresAt is t+d in epoch seconds rounded up, so a claim or commit never +// ends before it was asked to: TTL attributes are whole seconds. func expiresAt(t time.Time, d time.Duration) int64 { end := t.Add(d) sec := end.Unix() diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index e6276126..dc74d9d1 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -25,8 +25,8 @@ import ( // produce. type fakeDynamo struct { put func(context.Context, *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) - batch func(*dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) - del func(*dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) + batch func(context.Context, *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) + del func(context.Context, *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) describe func() (*dynamodb.DescribeTableOutput, error) ttl func() (*dynamodb.DescribeTimeToLiveOutput, error) } @@ -38,18 +38,24 @@ func (f *fakeDynamo) PutItem(ctx context.Context, in *dynamodb.PutItemInput, _ . return f.put(ctx, in) } -func (f *fakeDynamo) BatchWriteItem(_ context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { +func (f *fakeDynamo) BatchWriteItem(ctx context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { + if err := ctx.Err(); err != nil { + return nil, err + } if f.batch == nil { return &dynamodb.BatchWriteItemOutput{}, nil } - return f.batch(in) + return f.batch(ctx, in) } -func (f *fakeDynamo) DeleteItem(_ context.Context, in *dynamodb.DeleteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) { +func (f *fakeDynamo) DeleteItem(ctx context.Context, in *dynamodb.DeleteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) { + if err := ctx.Err(); err != nil { + return nil, err + } if f.del == nil { return &dynamodb.DeleteItemOutput{}, nil } - return f.del(in) + return f.del(ctx, in) } func (f *fakeDynamo) DescribeTable(context.Context, *dynamodb.DescribeTableInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTableOutput, error) { @@ -168,7 +174,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { } return &dynamodb.PutItemOutput{}, nil }, - del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) mu.Lock() defer mu.Unlock() @@ -179,7 +185,14 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { }) _, err := m.Reserve(t.Context(), keys("ok1", "dup", "bad", "ok2"), time.Minute) require.ErrorIs(t, err, ErrUnavailable) - assert.ElementsMatch(t, []string{"ok1", "bad", "ok2"}, released, "the failed put may have landed; the duplicate was never ours") + var sent []string + for id := range putTokens { + if id[len(id)-3:] != "dup" { + sent = append(sent, id[len(id)-3:]) + } + } + assert.ElementsMatch(t, sent, released, "every sent put but the duplicate, which was never ours") + assert.Contains(t, released, "bad", "the failed put may have landed") } func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { @@ -188,7 +201,7 @@ func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { var mu sync.Mutex written := map[string]int{} heldBack := map[string]bool{} - _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + _, m := openFake(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { calls.Add(1) reqs := in.RequestItems["dedupe"] assert.LessOrEqual(t, len(reqs), batchWriteMax) @@ -228,7 +241,7 @@ func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { t.Parallel() - _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + _, m := openFake(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { return &dynamodb.BatchWriteItemOutput{UnprocessedItems: in.RequestItems}, nil }}) err := m.Commit(t.Context(), []Claim{{Key: keys("a")[0], Status: Claimed, Token: "t"}}, 0) @@ -237,7 +250,7 @@ func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { t.Parallel() - _, m := openFake(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + _, m := openFake(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { assert.Equal(t, condRelease, aws.ToString(in.ConditionExpression)) id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) if id[len(id)-1] == 'x' { @@ -370,7 +383,7 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { t.Parallel() var mu sync.Mutex written := 0 - _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + _, m := openFakeWith(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { reqs := in.RequestItems["dedupe"] if idOf(reqs[0].PutRequest.Item[attrKey]) == "00" { return nil, &types.InternalServerError{} @@ -379,7 +392,7 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { written += len(reqs) mu.Unlock() return &dynamodb.BatchWriteItemOutput{}, nil - }}) + }}, DynamoConfig{Table: "dedupe", ReserveConcurrency: 1}) var claims []Claim for i := range 3 * batchWriteMax { claims = append(claims, Claim{Key: keys(fmt.Sprintf("%02d", i))[0], Status: Claimed, Token: "t"}) @@ -391,7 +404,7 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { t.Parallel() var deletes atomic.Int64 - _, m := openFakeWith(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + _, m := openFakeWith(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { deletes.Add(1) if idOf(in.Key[attrKey]) == "k0" { return nil, &types.ProvisionedThroughputExceededException{} @@ -424,7 +437,7 @@ func TestDynamo_FailedMultiKeyReserve(t *testing.T) { <-ctx.Done() return nil, ctx.Err() }, - del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { mu.Lock() released = append(released, idOf(in.Key[attrKey])) mu.Unlock() From 9824663384c5c2468583be4fb5f427bb3af50b86 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:26:22 -0400 Subject: [PATCH 026/108] docs(ingest): Release gives back definite failures only; tighten wording Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- 5 files changed, 5 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f0dab22e..195038a7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index dd84c32e..ca5f9b5c 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -272,7 +272,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | -| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now (not open, or a remote backend throttled or unreachable); `Retry-After: 5`. Nothing was published, so the retry is safe | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store is not open (for example, it failed to open on a reload); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue remembers the key for two minutes after the first publish, so a retry inside that window stores no second copy (the SDK's, after the 30-second `Retry-After`, lands inside it); a later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e50b269f..e0812d7d 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -122,7 +122,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `dedupe/` — Deduplication (Optional) -- **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. +- **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose records were definitely not published (a refused or never-sent publish; one whose outcome is unknown is left to lapse instead). A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 5f84c624..280bf458 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The window covers a prompt retry only while the lease is shorter than it. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. ## Check your storage before you trust it diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index d10be43e..1e898fdc 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,7 +186,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. From 99a5eee4b9ab7dd9f5fd46903581a298e4ace48d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:49:44 -0400 Subject: [PATCH 027/108] docs(ingest): say the Pebble benchmark stubs the queue; finish the rename Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/durability.md | 2 +- internal/api/ingest.go | 4 ++-- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 195038a7..87807233 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 280bf458..5a736b67 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -60,7 +60,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) ## Deduplication: one more fsync per window -With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index acf8bec8..6f129e34 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -505,7 +505,7 @@ func writeMaxBytesError(w http.ResponseWriter, err error, limit int64) bool { // // Evaluated here rather than per record because the condition is a property of // (table, role, policy) and is identical for every record in the request — the -// same reasoning as the !resolved abort in processRecord. Doing it per record +// same reasoning as the !resolved abort in prepareRecord. Doing it per record // would emit one ERROR line per record for a single mis-wired policy, which on // a 16 MiB body of small records is ~1.2M lines. The reject is still returned // per record, so a batch reports each record's own cause: one that SUPPLIES the @@ -518,7 +518,7 @@ func (h *IngestHandler) policyCheckGuard( ) *recordReject { checks, resolved := perms.CheckClauses() if !resolved { - return nil // the !resolved abort in processRecord owns this case + return nil // the !resolved abort in prepareRecord owns this case } // Sorted, and every offender — not the first one a map range happens to From ced118c0f0d2e697b13fda9d892f75c21189bc8c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 03:41:22 -0400 Subject: [PATCH 028/108] feat(app): choose the DynamoDB dedupe backend at boot dedupe.backend: dynamodb selects the shared table (F3's dedupe.Dynamo), configured by a dedupe.dynamodb block; dedupe.lease and dedupe.reserve_concurrency join the dedupe boot block. Boot checks the table, creating it first only with create_table on dynamodb-local; a failed check refuses a flat boot and fails switched-on tenants closed over a nested directory until a reload's check passes (Factory.Gated). The lease is capped at the embedded queue's 2m duplicate window. Part of #613 (PR F5). Co-Authored-By: Claude Opus 5.5 (1M context) --- AGENTS.md | 4 +- CHANGELOG.md | 3 +- config.yaml | 16 +- docs/src/content/docs/api.md | 4 +- docs/src/content/docs/architecture.md | 8 +- docs/src/content/docs/configuration.mdx | 51 +++++- docs/src/content/docs/deployment.md | 20 ++- docs/src/content/docs/settings-directory.mdx | 6 +- internal/app/app.go | 2 +- internal/app/dedupe_dynamodb_test.go | 164 +++++++++++++++++ internal/app/wire.go | 78 +++++++- internal/config/backends.go | 87 ++++++++- internal/config/backends_test.go | 143 ++++++++++++++- internal/dedupe/stores.go | 19 ++ internal/dedupe/stores_test.go | 19 ++ tests/integration/dedupe_dynamodb_app_test.go | 169 ++++++++++++++++++ 16 files changed, 758 insertions(+), 35 deletions(-) create mode 100644 internal/app/dedupe_dynamodb_test.go create mode 100644 tests/integration/dedupe_dynamodb_app_test.go diff --git a/AGENTS.md b/AGENTS.md index 8239899d..790a728c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,8 +34,8 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; built and conformance-tested against dynamodb-local but not yet selectable at boot), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (the in-process value by default; `dedupe.backend` also takes `dynamodb`, with its `dedupe.dynamodb` sub-block; `coord.backend` reserved) — boot is the validator, there is no dry run +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; conformance-tested against dynamodb-local, selected by `dedupe.backend: dynamodb`; boot checks the table and never creates it outside dynamodb-local), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index e60f59a6..f74c6cb5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/stores.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` Binary alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until a reload's check passes. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. diff --git a/config.yaml b/config.yaml index 5519b78c..7d4cbd45 100644 --- a/config.yaml +++ b/config.yaml @@ -43,12 +43,22 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none -# Each layer's implementation, chosen at boot. Only the in-process backend -# exists for each today, and it is the default. +# Each layer's implementation, chosen at boot. The in-process backend is +# each layer's default. mq: backend: embedded # NATS JetStream under /nats dedupe: - backend: pebble # Pebble under /pebble + backend: pebble # Pebble under /pebble; or dynamodb (below) + lease: 30s # how long a claimed id stays pending; at most 2m with the embedded mq + reserve_concurrency: 64 # parallel calls per request to a remote backend + # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain + # table: wavehouse-dedupe-prod + # region: "" # empty = AWS_REGION + # endpoint: "" # dynamodb-local only + # timeout: 250ms + # max_attempts: 3 + # retry_mode: standard # or adaptive + # create_table: false # dynamodb-local only; refused without endpoint coord: backend: local # reserved: nothing is elected yet diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index dc79ed6b..fdaf4696 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -275,7 +275,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | -| 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, [`dedupe.lease`](/configuration#dedupe), 30 seconds by default). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** @@ -387,7 +387,7 @@ A `200` is returned whenever the body was read and the records were processed | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. As for `publish failed`, the failing record's id is given back and the records before it keep theirs | -| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, [`dedupe.lease`](/configuration#dedupe), 30 seconds by default). The records before it were published | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4b10348d..0bb4030c 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with every reload checking again until it passes. It has no Pebble gauges. Both cases hand the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`). The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -126,10 +126,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, selected by `dedupe.backend: dynamodb`: every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. -- **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. +- **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `Factory.Gated(ready)` wraps a factory so a store opens only once `ready` returns nil, and fails closed until then (the DynamoDB wiring's table check). `internal/app` drives it from the registry's `AfterAdopt` hook. ### `discovery/` — Schema Discovery & Validation @@ -232,7 +232,7 @@ Client POST /v1/ingest?table={table} setting it to null is published un-deduped + logged/counted, or rejected under require_id); once the record is encoded, reserve (tenant, table, id): a duplicate is skipped, an id another request holds → 503 + Retry-After - (the 30s lease) + (dedupe.lease, 30s by default) → Publish to NATS JetStream (ingest.{tenant}.{table}) → Commit the reserved id; on a failed publish, release it instead → 200 OK returned immediately diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 193a6c21..3a7b8b12 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -39,16 +39,39 @@ This page is boot config only — what the platform operator owns (wiring, lifec ### Backends -Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. +Each layer's implementation is chosen once, at boot. The in-process backend is every layer's default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | | `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | -| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | +| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on; two processes do not share seen ids. `dynamodb`: one DynamoDB table that every tenant and every process shares, configured by [`dedupe.dynamodb`](#dynamodb-dedupe). | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. `dedupe.dynamodb` is the only one so far; any other, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. + +### Dedupe + +Whether a tenant dedupes, and on which field, are settings-directory keys ([Deduplication](/settings-directory#deduplication)). What is boot config is where the seen ids live and how a claim behaves. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`). | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most calls one request makes to a remote dedupe backend at once. `pebble` ignores it. | + +#### DynamoDB dedupe + +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (Binary) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and every reload checks again. The check runs whether or not any tenant has dedupe on. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(none)* | The shared table. Required. | +| `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`). | +| `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. | +| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. | +| `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. | +| `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | ### Server @@ -214,7 +237,17 @@ cache: l1_max_cost: 67108864 dedupe: - backend: pebble # in-process Pebble under /pebble + backend: pebble # in-process Pebble under /pebble; or dynamodb + lease: 30s # at most 2m with the embedded mq + reserve_concurrency: 64 + # dynamodb: # read only when backend is dynamodb + # table: wavehouse-dedupe-prod + # region: "" # empty = AWS_REGION + # endpoint: "" # dynamodb-local only + # timeout: 250ms + # max_attempts: 3 + # retry_mode: standard + # create_table: false # dynamodb-local only coord: backend: local # reserved: nothing is elected yet @@ -268,6 +301,16 @@ WH_MQ_BACKEND=embedded WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 WH_DEDUPE_BACKEND=pebble +WH_DEDUPE_LEASE=30s +WH_DEDUPE_RESERVE_CONCURRENCY=64 +# Read only with WH_DEDUPE_BACKEND=dynamodb: +# WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod +# WH_DEDUPE_DYNAMODB_REGION= +# WH_DEDUPE_DYNAMODB_ENDPOINT= +# WH_DEDUPE_DYNAMODB_TIMEOUT=250ms +# WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS=3 +# WH_DEDUPE_DYNAMODB_RETRY_MODE=standard +# WH_DEDUPE_DYNAMODB_CREATE_TABLE=false WH_COORD_BACKEND=local WH_AUTH_JWT_SECRET=change-me-in-production diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 9bea802e..5d0cb186 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -423,10 +423,6 @@ The dedupe key now carries the table as well as the tenant ([#222](https://githu ## A shared dedupe table on DynamoDB -:::note[Not selectable yet] -The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. A boot key to select it lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. -::: - Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend refuses to create a table unless it is pointed at a custom endpoint, so table creation only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). What the backend requires of the table: @@ -438,7 +434,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table: it refuses one whose key schema does not match, and logs a warning if TTL is off. An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: @@ -487,6 +483,20 @@ data "aws_iam_policy_document" "wavehouse_dedupe" { } ``` +Select it in the boot config, on every pod that should share seen ids (all the keys are in the [Configuration Reference](/configuration#dynamodb-dedupe)): + +```yaml +dedupe: + backend: dynamodb + dynamodb: + table: wavehouse-dedupe-prod + region: us-east-1 # or leave empty for AWS_REGION +``` + +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and each reload checks the table again. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. + +For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. + - **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. - **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index df099659..156421f7 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -183,10 +183,10 @@ What stays in boot config is only what cannot change under a running process — ## Deduplication -Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): +Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/app.go b/internal/app/app.go index 51935e7d..d8f18283 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -174,7 +174,7 @@ func New(ctx context.Context, opts Options) (app *App, err error) { return nil, err } a.wireDiscovery(ctx) - if err := a.wireDedupe(); err != nil { + if err := a.wireDedupe(ctx); err != nil { return nil, err } if err := a.wireMQ(ctx); err != nil { diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go new file mode 100644 index 00000000..ded483a9 --- /dev/null +++ b/internal/app/dedupe_dynamodb_test.go @@ -0,0 +1,164 @@ +package app + +import ( + "context" + "io" + "net/http" + "net/http/httptest" + "path/filepath" + "strings" + "sync" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/config" + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/dedupe/dedupetest" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// fakeDynamo answers the DynamoDB JSON protocol for one table, enough for +// boot's check, the dev create path, and a claim and its commit. Whether the +// table exists is the test's to switch. +type fakeDynamo struct { + mu sync.Mutex + exists bool + calls []string +} + +func (f *fakeDynamo) setExists(v bool) { + f.mu.Lock() + defer f.mu.Unlock() + f.exists = v +} + +func (f *fakeDynamo) called(op string) bool { + f.mu.Lock() + defer f.mu.Unlock() + for _, c := range f.calls { + if c == op { + return true + } + } + return false +} + +func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { + _, _ = io.Copy(io.Discard, r.Body) + _, op, _ := strings.Cut(r.Header.Get("X-Amz-Target"), ".") + f.mu.Lock() + f.calls = append(f.calls, op) + if op == "CreateTable" { + f.exists = true + } + exists := f.exists + f.mu.Unlock() + w.Header().Set("Content-Type", "application/x-amz-json-1.0") + if !exists { + w.WriteHeader(http.StatusBadRequest) + _, _ = io.WriteString(w, `{"__type":"com.amazonaws.dynamodb.v20120810#ResourceNotFoundException","message":"Requested resource not found"}`) + return + } + body := `{}` + switch op { + case "DescribeTable", "CreateTable": + body = `{"Table":{"TableName":"dedupe","TableStatus":"ACTIVE",` + + `"KeySchema":[{"AttributeName":"pk","KeyType":"HASH"}],` + + `"AttributeDefinitions":[{"AttributeName":"pk","AttributeType":"B"}]}}` + case "DescribeTimeToLive": + body = `{"TimeToLiveDescription":{"AttributeName":"ex","TimeToLiveStatus":"ENABLED"}}` + case "BatchWriteItem": + body = `{"UnprocessedItems":{}}` + } + _, _ = io.WriteString(w, body) +} + +// dynamoConfig points cfg's dedupe at a fake table, with credentials from the +// environment as the SDK's default chain reads them — and nothing from the +// developer's own AWS files. +func dynamoConfig(t *testing.T, cfg *config.Config, exists bool) *fakeDynamo { + t.Helper() + fake := &fakeDynamo{exists: exists} + srv := httptest.NewServer(fake) + t.Cleanup(srv.Close) + none := filepath.Join(t.TempDir(), "none") + for k, v := range map[string]string{ + "AWS_ACCESS_KEY_ID": "local", "AWS_SECRET_ACCESS_KEY": "local", "AWS_SESSION_TOKEN": "", + "AWS_PROFILE": "", "AWS_CONFIG_FILE": none, "AWS_SHARED_CREDENTIALS_FILE": none, + "AWS_EC2_METADATA_DISABLED": "true", + } { + t.Setenv(k, v) + } + cfg.Dedupe = config.Dedupe{Backend: config.DedupeDynamoDB, DynamoDB: config.DedupeDynamoDBConfig{ + Table: "dedupe", Region: "us-east-1", Endpoint: srv.URL, MaxAttempts: 1, + }} + return fake +} + +var dedupeOn = map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}} + +func TestNew_DynamoDBDedupe(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, true) + a := newApp(t, cfg, Options{}) + + assert.True(t, fake.called("DescribeTable"), "boot checks the table") + assert.False(t, fake.called("CreateTable"), "and never creates it without create_table") + store := a.dedup.For(tenant.Default) + require.True(t, store.Open()) + dup, err := dedupetest.Mark(t.Context(), store, eventKey) + require.NoError(t, err) + assert.False(t, dup) + assert.True(t, fake.called("PutItem"), "the claim went to the table") + assert.True(t, fake.called("BatchWriteItem"), "and so did its commit") + assert.Nil(t, a.dedupeStats, "no Pebble instance, so no Pebble gauges") + assert.NoDirExists(t, filepath.Join(cfg.DataDir, "pebble")) +} + +func TestNew_DynamoDBDedupeCreatesTheTableOnlyWhenAsked(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, false) + cfg.Dedupe.DynamoDB.CreateTable = true + a := newApp(t, cfg, Options{}) + assert.True(t, fake.called("CreateTable")) + assert.True(t, fake.called("UpdateTimeToLive")) + assert.True(t, a.dedup.For(tenant.Default).Open()) +} + +// A table that fails the check follows the registry's rule for the shape, +// as a Pebble instance that cannot open does. +func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { + t.Run("flat refuses boot", func(t *testing.T) { + for name, patch := range map[string]map[string]any{"dedupe on": dedupeOn, "dedupe off": nil} { + t.Run(name, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, patch)) + dynamoConfig(t, cfg, false) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "dedupe open") + require.ErrorContains(t, err, "ResourceNotFoundException") + }) + } + }) + t.Run("nested fails closed until a reload passes the check", func(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + a := newApp(t, cfg, Options{}) + + acme := a.dedup.For("acme") + assert.False(t, acme.Open()) + _, err := dedupetest.Mark(t.Context(), acme, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on, table missing: ingest fails closed") + _, err = dedupetest.Mark(t.Context(), a.dedup.For("globex"), eventKey) + require.ErrorIs(t, err, dedupe.ErrDisabled) + + fake.setExists(true) + a.tenants.Reload("test") + assert.True(t, acme.Open(), "the reload checked again and opened the store") + _, err = dedupetest.Mark(context.Background(), acme, eventKey) + require.NoError(t, err) + }) +} diff --git a/internal/app/wire.go b/internal/app/wire.go index 02164497..426db887 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -461,10 +461,12 @@ func (a *App) wireDiscovery(ctx context.Context) { // wireDedupe builds the dedupe stores — the one place the implementation is // chosen. -func (a *App) wireDedupe() error { +func (a *App) wireDedupe(ctx context.Context) error { switch b := a.cfg.Dedupe.Backend; b { case config.DedupePebble: return a.wirePebbleDedupe() + case config.DedupeDynamoDB: + return a.wireDynamoDedupe(ctx) default: return unreachableBackend("dedupe.backend", b) } @@ -529,6 +531,79 @@ func (a *App) wirePebbleDedupe() error { return nil } +// errDynamoUnchecked is a store's open before the first table check has run. +var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") + +// wireDynamoDedupe builds the dedupe stores over one DynamoDB table that +// every tenant and every process shares (dedupe.Dynamo), so a tenant's store +// opens for free once the table has passed its check. Boot checks it (after +// creating it, with create_table on dynamodb-local) whether or not any tenant +// has dedupe on, and never creates it otherwise. A table that fails the check +// follows the registry's rule for the shape, as Pebble's instance does: a +// flat directory refuses boot; a nested one boots with every switched-on +// store closed, so its ingest fails closed, and each reload checks again. +func (a *App) wireDynamoDedupe(ctx context.Context) error { + c := a.cfg.Dedupe.DynamoDB + d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ + Table: c.Table, Region: c.Region, Endpoint: c.Endpoint, + Timeout: c.Timeout, MaxAttempts: c.MaxAttempts, RetryMode: c.RetryMode, + ReserveConcurrency: a.cfg.Dedupe.ReserveConcurrency, + }) + if err != nil { + return err + } + var mu sync.Mutex + state := errDynamoUnchecked // nil once the table has passed + check := func(ctx context.Context) error { + mu.Lock() + defer mu.Unlock() + if state == nil { + return nil + } + if c.CreateTable { + if state = d.CreateTable(ctx); state != nil { + return state + } + } + state = d.Check(ctx) + return state + } + ready := func() error { + mu.Lock() + defer mu.Unlock() + return state + } + stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) + a.dedup = stores + a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) + reconcile := func(ctx context.Context) error { + if err := stores.Retain(a.served); err != nil { + slog.Error("dedupe store close failed", "error", err) + } + checkErr := check(ctx) + if checkErr != nil { + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + "table", c.Table, "error", checkErr) + } + for id, store := range a.tenants.All() { + m := stores.For(id) + enabled := store.DedupeEnabled() + wasOpen := m.Open() + // The one failure an open has is the check's, logged above. + _ = m.Apply(enabled) + if m.Open() != wasOpen { + slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) + } + } + return checkErr + } + a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) + if err := reconcile(ctx); err != nil && !a.tenants.Nested() { + return fmt.Errorf("dedupe open: %w", err) + } + return nil +} + // wireMQ starts the MQ — the one place the implementation is chosen; // everything after it sees mq.Broker. func (a *App) wireMQ(ctx context.Context) error { @@ -867,6 +942,7 @@ func (a *App) wireHTTP(authMW func(http.Handler) http.Handler) { ingestHandler.PolicySource = (*settings.Store).Policy ingestHandler.Dedup = func(s *settings.Store) dedupe.Deduplicator { return a.dedup.For(s.Tenant()) } ingestHandler.DedupeSettings = (*settings.Store).DedupeFor + ingestHandler.DedupeLease = a.cfg.Dedupe.Lease // Readiness pings every open pool at once and is ready at the first // answer: one tenant's ClickHouse outage is not the process's. diff --git a/internal/config/backends.go b/internal/config/backends.go index f2ab9330..9b0cb57a 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -1,9 +1,11 @@ package config import ( + "errors" "fmt" "slices" "strings" + "time" ) // Each layer's implementation is chosen here, once, at boot: `.backend` @@ -55,20 +57,80 @@ func (c Cache) validate() error { // DedupeBackend names where ingest dedupe keeps the ids it has seen. type DedupeBackend string -// DedupePebble is the Pebble instance inside this process, under -// /pebble, opened while any tenant has dedupe on. -const DedupePebble DedupeBackend = "pebble" +const ( + // DedupePebble is the Pebble instance inside this process, under + // /pebble, opened while any tenant has dedupe on. Seen ids are + // per process. + DedupePebble DedupeBackend = "pebble" + // DedupeDynamoDB is one DynamoDB table every tenant and every process + // shares, configured by dedupe.dynamodb. + DedupeDynamoDB DedupeBackend = "dynamodb" +) -var dedupeBackends = []DedupeBackend{DedupePebble} +var dedupeBackends = []DedupeBackend{DedupePebble, DedupeDynamoDB} -// Dedupe selects the dedupe store. Whether a tenant dedupes, and on which -// field, are settings-directory keys, not this block's. +// Dedupe selects the dedupe store. Whether a tenant dedupes, on which field, +// and for how long are settings-directory keys, not this block's. type Dedupe struct { Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND" env-default:"pebble"` + // Lease is how long a claimed id stays pending while its record is + // published; a claim its request never settles lapses after it. + Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE" env-default:"30s"` + // ReserveConcurrency bounds the parallel calls one request makes to a + // remote backend. Pebble ignores it. + ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY" env-default:"64"` + DynamoDB DedupeDynamoDBConfig `yaml:"dynamodb"` +} + +// DedupeDynamoDBConfig is the dynamodb backend's block, read only when it is +// selected. Credentials are the AWS SDK's default chain (EKS Pod Identity, +// IRSA, AWS_* variables), never keys here. +type DedupeDynamoDBConfig struct { + // Table is the shared table; WaveHouse never creates it outside + // dynamodb-local. Required. + Table string `yaml:"table" env:"WH_DEDUPE_DYNAMODB_TABLE"` + // Region overrides the SDK chain's (AWS_REGION). + Region string `yaml:"region" env:"WH_DEDUPE_DYNAMODB_REGION"` + // Endpoint points the client at dynamodb-local. + Endpoint string `yaml:"endpoint" env:"WH_DEDUPE_DYNAMODB_ENDPOINT"` + Timeout time.Duration `yaml:"timeout" env:"WH_DEDUPE_DYNAMODB_TIMEOUT" env-default:"250ms"` + MaxAttempts int `yaml:"max_attempts" env:"WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS" env-default:"3"` + RetryMode string `yaml:"retry_mode" env:"WH_DEDUPE_DYNAMODB_RETRY_MODE" env-default:"standard"` + // CreateTable creates the table at boot if it is missing. Development + // only: refused unless Endpoint is set. + CreateTable bool `yaml:"create_table" env:"WH_DEDUPE_DYNAMODB_CREATE_TABLE" env-default:"false"` } func (d Dedupe) validate() error { - return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) + if err := checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends); err != nil { + return err + } + if d.Lease < 0 { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) must be >= 0, got %s", d.Lease) + } + if d.ReserveConcurrency < 0 { + return fmt.Errorf("dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be >= 0, got %d", d.ReserveConcurrency) + } + if d.Backend == DedupeDynamoDB { + return d.DynamoDB.validate() + } + return nil +} + +func (d DedupeDynamoDBConfig) validate() error { + switch { + case strings.TrimSpace(d.Table) == "": + return errors.New("dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required when dedupe.backend is dynamodb") + case d.Timeout < 0: + return fmt.Errorf("dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be >= 0, got %s", d.Timeout) + case d.MaxAttempts < 0: + return fmt.Errorf("dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be >= 0, got %d", d.MaxAttempts) + case d.RetryMode != "" && d.RetryMode != "standard" && d.RetryMode != "adaptive": + return fmt.Errorf("dedupe.dynamodb.retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) %q: want standard or adaptive", d.RetryMode) + case d.CreateTable && d.Endpoint == "": + return errors.New("dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only: set dedupe.dynamodb.endpoint, or create the table with your infrastructure code") + } + return nil } // CoordBackend names where leases for singleton work (the sweeper) are held. @@ -104,13 +166,22 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) } -// validateBackends checks every layer's backend and its sub-block. +// embeddedDuplicateWindow mirrors mq.EmbeddedDuplicateWindow, the embedded +// ingest stream's duplicate window (#613 F2). A lease longer than it would let +// the republish of a publish whose outcome was unknown land twice. +const embeddedDuplicateWindow = 2 * time.Minute + +// validateBackends checks every layer's backend and its sub-block, then the +// rules that span two layers. func (c *Config) validateBackends() error { for _, check := range []func() error{c.MQ.validate, c.Cache.validate, c.Dedupe.validate, c.Coord.validate} { if err := check(); err != nil { return err } } + if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > embeddedDuplicateWindow { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s exceeds the embedded mq's %s duplicate window: a claim must lapse before the queue forgets the publish it guards", c.Dedupe.Lease, embeddedDuplicateWindow) + } return nil } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 0e70dc7a..803b48d7 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -4,6 +4,7 @@ import ( "os" "path/filepath" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -28,6 +29,10 @@ func TestLoad_BackendDefaults(t *testing.T) { assert.Equal(t, MQEmbedded, cfg.MQ.Backend) assert.Equal(t, CacheLocal, cfg.Cache.Backend) assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, Dedupe{ + Backend: DedupePebble, Lease: 30 * time.Second, ReserveConcurrency: 64, + DynamoDB: DedupeDynamoDBConfig{Timeout: 250 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard"}, + }, cfg.Dedupe) assert.Equal(t, CoordLocal, cfg.Coord.Backend) assert.False(t, cfg.Distributed()) assert.True(t, cfg.NeedsDataDir()) @@ -112,7 +117,7 @@ func TestValidate_UnknownBackend(t *testing.T) { }{ {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, - {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, + {"dedupe", func(c *Config) { c.Dedupe.Backend = "redis" }, `dedupe.backend (WH_DEDUPE_BACKEND) "redis" is not a backend this build has; valid: pebble, dynamodb`}, {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, // The zero value, which a Config built without Load carries. {"empty", func(c *Config) { c.MQ.Backend = "" }, `mq.backend (WH_MQ_BACKEND) "" is not a backend`}, @@ -159,3 +164,139 @@ func TestNeedsDataDir(t *testing.T) { cfg.MQ.Backend = MQEmbedded assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") } + +func TestLoad_DedupeDynamoDBFromEnv(t *testing.T) { + for k, v := range map[string]string{ + "WH_DEDUPE_BACKEND": "dynamodb", + "WH_DEDUPE_LEASE": "45s", + "WH_DEDUPE_RESERVE_CONCURRENCY": "16", + "WH_DEDUPE_DYNAMODB_TABLE": "wavehouse-dedupe-dev", + "WH_DEDUPE_DYNAMODB_REGION": "us-east-2", + "WH_DEDUPE_DYNAMODB_ENDPOINT": "http://localhost:8000", + "WH_DEDUPE_DYNAMODB_TIMEOUT": "1s", + "WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS": "5", + "WH_DEDUPE_DYNAMODB_RETRY_MODE": "adaptive", + "WH_DEDUPE_DYNAMODB_CREATE_TABLE": "true", + } { + t.Setenv(k, v) + } + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, Dedupe{ + Backend: DedupeDynamoDB, Lease: 45 * time.Second, ReserveConcurrency: 16, + DynamoDB: DedupeDynamoDBConfig{ + Table: "wavehouse-dedupe-dev", Region: "us-east-2", Endpoint: "http://localhost:8000", + Timeout: time.Second, MaxAttempts: 5, RetryMode: "adaptive", CreateTable: true, + }, + }, cfg.Dedupe) + assert.True(t, cfg.NeedsDataDir(), "the embedded mq still keeps state under data_dir") +} + +func TestLoad_DedupeDynamoDBFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +dedupe: + backend: dynamodb + lease: 20s + dynamodb: + table: wavehouse-dedupe-prod + timeout: 400ms +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, DedupeDynamoDB, cfg.Dedupe.Backend) + assert.Equal(t, 20*time.Second, cfg.Dedupe.Lease) + assert.Equal(t, 64, cfg.Dedupe.ReserveConcurrency) + assert.Equal(t, DedupeDynamoDBConfig{ + Table: "wavehouse-dedupe-prod", Timeout: 400 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard", + }, cfg.Dedupe.DynamoDB) +} + +func TestLoad_DedupeDynamoDBRefusesUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +dedupe: + backend: dynamodb + dynamodb: + table: t + access_key_id: AKIA + redis: + addr: localhost:6379 +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "dedupe.dynamodb.access_key_id, dedupe.redis") +} + +func TestUnboundEnv_KnowsTheDedupeVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_DEDUPE_LEASE=30s", "WH_DEDUPE_RESERVE_CONCURRENCY=64", + "WH_DEDUPE_DYNAMODB_TABLE=t", "WH_DEDUPE_DYNAMODB_REGION=us-east-1", + "WH_DEDUPE_DYNAMODB_ENDPOINT=http://localhost:8000", "WH_DEDUPE_DYNAMODB_TIMEOUT=250ms", + "WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS=3", "WH_DEDUPE_DYNAMODB_RETRY_MODE=standard", + "WH_DEDUPE_DYNAMODB_CREATE_TABLE=false", + })) +} + +func TestValidate_Dedupe(t *testing.T) { + t.Parallel() + dynamo := func(c *Config) { + c.Dedupe.Backend = DedupeDynamoDB + c.Dedupe.DynamoDB = DedupeDynamoDBConfig{Table: "t", Timeout: time.Second, MaxAttempts: 3, RetryMode: "standard"} + } + cases := []struct { + name string + set func(*Config) + want string // "" = valid + }{ + {"dynamodb", dynamo, ""}, + {"zero values read as the defaults", func(c *Config) { + c.Dedupe.Backend = DedupeDynamoDB + c.Dedupe.DynamoDB = DedupeDynamoDBConfig{Table: "t"} + }, ""}, + {"create_table with an endpoint", func(c *Config) { + dynamo(c) + c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true + }, ""}, + {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB.CreateTable = true }, ""}, + {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, ""}, + {"create_table without an endpoint", func(c *Config) { + dynamo(c) + c.Dedupe.DynamoDB.CreateTable = true + }, "dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only"}, + {"no table", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Table = " " }, "dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required"}, + {"retry mode", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.RetryMode = "legacy" }, `retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) "legacy"`}, + {"negative timeout", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Timeout = -time.Second }, "dedupe.dynamodb.timeout"}, + {"negative attempts", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.MaxAttempts = -1 }, "dedupe.dynamodb.max_attempts"}, + {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be >= 0"}, + {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, + {"lease past the duplicate window", func(c *Config) { c.Dedupe.Lease = 3 * time.Minute }, "exceeds the embedded mq's 2m0s duplicate window"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + tc.set(&cfg) + err := cfg.Validate() + if tc.want == "" { + require.NoError(t, err) + return + } + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +func TestNeedsDataDir_DynamoDBDedupe(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + cfg.Dedupe.Backend = DedupeDynamoDB + assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") + cfg.MQ.Backend = "shared" + assert.False(t, cfg.NeedsDataDir(), "neither a shared mq nor dynamodb dedupe keeps state under data_dir") + assert.Len(t, cfg.Warnings(), 1, "only the local cache warning: dynamodb dedupe is shared") +} diff --git a/internal/dedupe/stores.go b/internal/dedupe/stores.go index 614912fe..5bb39366 100644 --- a/internal/dedupe/stores.go +++ b/internal/dedupe/stores.go @@ -16,6 +16,25 @@ import ( // nothing that holds the Stores changes with it. type Factory func(id tenant.ID) *Managed +// Gated returns a Factory whose stores open only once ready returns nil, its +// error being the open's: a store switched on meanwhile stays closed and +// fails closed (ErrUnavailable) until an Apply finds the backend ready. For a +// backend whose tenant opens are free but whose shared resource (a remote +// table) is checked once. +func (f Factory) Gated(ready func() error) Factory { + return func(id tenant.ID) *Managed { + m := f(id) + open := m.open + m.open = func() (Deduplicator, error) { + if err := ready(); err != nil { + return nil, err + } + return open() + } + return m + } +} + // Stores is one Managed store per tenant (#583 story 7), each following its // own tenant's dedupe.enabled through Apply. A store is built on first use // and forgotten by Retain once its tenant is no longer served; its seen ids diff --git a/internal/dedupe/stores_test.go b/internal/dedupe/stores_test.go index ea2ae065..03e6ecc6 100644 --- a/internal/dedupe/stores_test.go +++ b/internal/dedupe/stores_test.go @@ -2,6 +2,7 @@ package dedupe import ( "context" + "errors" "testing" "time" @@ -134,3 +135,21 @@ func TestStores_CloseClosesEveryStore(t *testing.T) { assert.False(t, e.Open(), "the instance closes with the last store") require.NoError(t, s.Close(), "closing again is a no-op") } + +func TestFactory_GatedOpensOnlyOnceReady(t *testing.T) { + t.Parallel() + notReady := errors.New("table missing") + ready := notReady + gated := NewStores(Factory(NewEmbedded(t.TempDir()).Tenant).Gated(func() error { return ready })) + t.Cleanup(func() { _ = gated.Close() }) + acme := gated.For("acme") + + require.ErrorIs(t, acme.Apply(true), notReady) + assert.False(t, acme.Open()) + _, err := mark(context.Background(), acme, "e1") + require.ErrorIs(t, err, ErrUnavailable, "switched on but not ready: fails closed, never open") + + ready = nil + require.NoError(t, acme.Apply(true), "the next apply finds it ready") + assert.True(t, acme.Open()) +} diff --git a/tests/integration/dedupe_dynamodb_app_test.go b/tests/integration/dedupe_dynamodb_app_test.go new file mode 100644 index 00000000..f94ccb4f --- /dev/null +++ b/tests/integration/dedupe_dynamodb_app_test.go @@ -0,0 +1,169 @@ +//go:build integration + +package tests + +import ( + "context" + "encoding/json" + "fmt" + "io" + "net" + "net/http" + "net/url" + "path/filepath" + "strings" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/app" + "github.com/Wave-RF/WaveHouse/internal/config" + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// TestDynamoDBDedupe_TwoInstancesShareSeenIDs boots two apps the way two +// pods run — each its own data_dir, embedded queue and ingest worker — with +// dedupe.backend dynamodb over one table on dynamodb-local, and checks an id +// ingested through either is a duplicate through the other, that ClickHouse +// holds each id once, and that dedupe.lease reaches ingest as the in-flight +// answer's Retry-After. +func TestDynamoDBDedupe_TwoInstancesShareSeenIDs(t *testing.T) { + e := env(t) + ctx := context.Background() + // The SDK's default chain, as in production; never the developer's files. + none := filepath.Join(t.TempDir(), "none") + for k, v := range map[string]string{ + "AWS_ACCESS_KEY_ID": "local", "AWS_SECRET_ACCESS_KEY": "local", "AWS_SESSION_TOKEN": "", + "AWS_PROFILE": "", "AWS_CONFIG_FILE": none, "AWS_SHARED_CREDENTIALS_FILE": none, + "AWS_EC2_METADATA_DISABLED": "true", + } { + t.Setenv(k, v) + } + + chTable := createTable(t, "event_id String, n UInt32", "ORDER BY event_id") + ddbTable := newDynamoTable() + const lease = 7 * time.Second + + boot := func(name string) string { + t.Helper() + files, err := tenantSettings(e.ch, testCHDatabase) + require.NoError(t, err) + var doc map[string]json.RawMessage + require.NoError(t, json.Unmarshal(files[settings.FileConfig], &doc)) + doc["dedupe"] = json.RawMessage(`{"enabled": true, "id_field": "event_id", "require_id": true, "tables": {}}`) + files[settings.FileConfig], err = json.Marshal(doc) + require.NoError(t, err) + dir := filepath.Join(t.TempDir(), name) + require.NoError(t, writeSettingsFiles(dir, files)) + + var lc net.ListenConfig + ln, err := lc.Listen(ctx, "tcp", "127.0.0.1:0") + require.NoError(t, err) + cfg := &config.Config{ + DataDir: t.TempDir(), + Server: config.Server{ShutdownTimeout: 10}, + ClickHouse: config.ClickHouse{Password: testCHPassword}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupeDynamoDB, Lease: lease, DynamoDB: config.DedupeDynamoDBConfig{ + Table: ddbTable, Region: "us-east-1", Endpoint: e.dynamoEndpoint, + // dynamodb-local under a parallel suite is slower than the real thing. + Timeout: 5 * time.Second, CreateTable: true, + }}, + Coord: config.Coord{Backend: config.CoordLocal}, + Settings: config.Settings{Dir: dir}, + } + a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) + require.NoError(t, err) + runCtx, stop := context.WithCancel(ctx) + runDone := make(chan error, 1) + go func() { runDone <- a.Run(runCtx) }() + t.Cleanup(func() { + stop() + assert.NoError(t, <-runDone) + closeCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second) + defer cancel() + assert.NoError(t, a.Close(closeCtx)) + }) + baseURL := "http://" + ln.Addr().String() + require.NoError(t, waitForLive(ctx, baseURL, 30*time.Second)) + return baseURL + } + // Both create the table: the second finds it and leaves it as it is. + podA, podB := boot("a"), boot("b") + + ingest := func(baseURL, id string, n int) (int, string, http.Header) { + t.Helper() + body := fmt.Sprintf(`{"event_id": %q, "n": %d}`, id, n) + req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+"/v1/ingest?table="+url.QueryEscape(chTable), strings.NewReader(body)) + require.NoError(t, err) + req.Header.Set("Content-Type", "application/json") + resp, err := http.DefaultClient.Do(req) + require.NoError(t, err) + defer func() { _ = resp.Body.Close() }() + b, err := io.ReadAll(resp.Body) + require.NoError(t, err) + return resp.StatusCode, strings.TrimSpace(string(b)), resp.Header + } + accepted := func(baseURL, id string, n int) { + t.Helper() + status, body, _ := ingest(baseURL, id, n) + require.Equal(t, http.StatusOK, status, body) + require.JSONEq(t, `{"ok": true}`, body) + } + duplicate := func(baseURL, id string, n int) { + t.Helper() + status, body, _ := ingest(baseURL, id, n) + require.Equal(t, http.StatusOK, status, body) + require.JSONEq(t, `{"duplicate": true}`, body) + } + + accepted(podA, "e1", 1) + duplicate(podB, "e1", 2) + accepted(podB, "e2", 3) + duplicate(podA, "e2", 4) + duplicate(podA, "e1", 5) + + // A claim another process holds is in flight on both pods, for as long + // as the configured lease says. + peer := dynamoClient(t, ddbTable, dedupe.DynamoConfig{}).Tenant(tenant.Default) + require.NoError(t, peer.Apply(true)) + claims, err := peer.Reserve(ctx, []dedupe.Key{{Table: chTable, ID: "e3"}}, time.Minute) + require.NoError(t, err) + require.Equal(t, dedupe.Claimed, claims[0].Status) + for _, pod := range []string{podA, podB} { + status, body, header := ingest(pod, "e3", 6) + require.Equal(t, http.StatusServiceUnavailable, status, body) + assert.Equal(t, "7", header.Get("Retry-After"), "dedupe.lease, in seconds") + } + require.NoError(t, peer.Release(ctx, claims)) + accepted(podB, "e3", 7) + duplicate(podA, "e3", 8) + + // Each pod's worker wrote only what its pod accepted: each id once. + type row struct { + ID string + N uint32 + } + want := []row{{"e1", 1}, {"e2", 3}, {"e3", 7}} + require.Eventually(t, func() bool { + rows, err := e.chConn.Query(ctx, fmt.Sprintf("SELECT event_id, n FROM %s ORDER BY event_id", chTable)) + if err != nil { + return false + } + defer func() { _ = rows.Close() }() + var got []row + for rows.Next() { + var r row + if rows.Scan(&r.ID, &r.N) != nil { + return false + } + got = append(got, r) + } + return assert.ObjectsAreEqual(want, got) + }, 30*time.Second, 500*time.Millisecond, "ClickHouse holds each id once, from the pod that accepted it") +} From b83e67075a85017e31ae7d31fabd1b1bf939d35e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 03:49:16 -0400 Subject: [PATCH 029/108] feat(dedupe): retention per tenant and table, and an expiry sweep dedupe.retention (required, "0" = forever) and its per-table override are read per record and passed to Commit. Validation refuses a finite retention below the queue's two-minute duplicate window. The embedded Pebble store deletes expired keys and the version-0 keys in an hourly background sweep, counted by wavehouse_dedupe_swept_keys_total. Part of #613. Refs #220. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 3 +- cmd/wavehouse/validate_test.go | 2 +- config.yaml | 2 +- deployments/compose/settings/config.json | 1 + docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 6 +- docs/src/content/docs/durability.md | 4 +- docs/src/content/docs/settings-directory.mdx | 9 +- internal/api/ingest.go | 66 +++++--- internal/api/ingest_retention_test.go | 104 ++++++++++++ internal/api/ingest_test.go | 40 +++-- internal/api/ingest_window_test.go | 4 +- internal/api/settings_test.go | 2 +- internal/app/app_test.go | 14 +- internal/dedupe/embedded.go | 65 ++++++-- internal/dedupe/sweep.go | 151 ++++++++++++++++++ internal/dedupe/sweep_test.go | 158 +++++++++++++++++++ internal/settings/registry_test.go | 4 +- internal/settings/seed/config.json | 1 + internal/settings/settings.go | 18 ++- internal/settings/store.go | 31 +++- internal/settings/store_test.go | 21 ++- internal/settings/validate.go | 30 +++- internal/settings/validate_test.go | 39 ++++- internal/testutil/mocks.go | 13 +- tests/e2e/fixtures/settings/config.json | 1 + 27 files changed, 693 insertions(+), 102 deletions(-) create mode 100644 internal/api/ingest_retention_test.go create mode 100644 internal/dedupe/sweep.go create mode 100644 internal/dedupe/sweep_test.go diff --git a/AGENTS.md b/AGENTS.md index ffbd3c16..b26db3e3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, committed ids stored with their expiry and deleted by an hourly background sweep along with the version-0 keys from before the table joined the key, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal; `WithIdempotencyKey` makes a republish inside the queue's duplicate window a no-op), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` @@ -58,7 +58,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. 6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format`, or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. -8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase, one call per phase per window of up to 256 records — `Reserve` → publish (under the id's idempotency key) → `Commit`, or `Release` when the publish definitely failed, while one whose outcome is unknown is left to lapse; a store that cannot answer is a `503` + `Retry-After`; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. +8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase, one call per phase per window of up to 256 records — `Reserve` → publish (under the id's idempotency key) → `Commit`, or `Release` when the publish definitely failed, while one whose outcome is unknown is left to lapse; a store that cannot answer is a `503` + `Retry-After`; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key and `dedupe.retention` how long a committed id stays a duplicate (`"0"` = forever, else at least the queue's two-minute duplicate window), both overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 87807233..f47f81ba 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,6 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It deletes 1,024 keys per chunk without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed @@ -78,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, deleted by the retention sweep below ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/cmd/wavehouse/validate_test.go b/cmd/wavehouse/validate_test.go index e2e18e6c..5598cf65 100644 --- a/cmd/wavehouse/validate_test.go +++ b/cmd/wavehouse/validate_test.go @@ -19,7 +19,7 @@ func writeSettingsDir(t *testing.T, policies string) string { "roles.json": `{"roles": ["public"]}`, "policies.json": policies, "pipes.json": `{}`, - "config.json": `{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 10000, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, + "config.json": `{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 10000, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, } for name, content := range files { require.NoError(t, os.WriteFile(filepath.Join(dir, name), []byte(content), 0o600)) diff --git a/config.yaml b/config.yaml index 53a43502..35812417 100644 --- a/config.yaml +++ b/config.yaml @@ -63,7 +63,7 @@ auth: # clickhouse wiring (addr, http_port, http_scheme, database, username, # query_timeout, tls, headers, max_open_conns, max_idle_conns), auth # (jwks_url, role_claim), dedupe (enabled/id_field/ -# require_id + per-table overrides), dlq.enabled (+ per table), +# require_id/retention + per-table overrides), dlq.enabled (+ per table), # query.default_max_rows / timestamp_bucket_seconds, # schema.refresh_interval, stream keepalive_interval / keepalive_buckets / # gap_window_minutes, mq.max_bytes_gb, cors.allowed_origins — and every key diff --git a/deployments/compose/settings/config.json b/deployments/compose/settings/config.json index 030d76cc..6b33f55c 100644 --- a/deployments/compose/settings/config.json +++ b/deployments/compose/settings/config.json @@ -26,6 +26,7 @@ "enabled": false, "id_field": "event_id", "require_id": false, + "retention": "0", "tables": {} }, "dlq": { diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e0812d7d..29b6b284 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -124,7 +124,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose records were definitely not published (a refused or never-sent publish; one whose outcome is unknown is left to lapse instead). A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. -- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync, each value carrying its expiry (`0` = never), which `Reserve` honors on read. A background sweep (`sweep.go`), started when the instance opens and stopped before it closes, deletes expired keys and the version-0 keys from before the table joined the key ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)): a minute after opening, then hourly, 1,024 keys per chunk, holding a lock `Commit` also takes, so a key re-committed after the sweep read it is never deleted; `wavehouse_dedupe_swept_keys_total{reason}` counts what it deletes. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index bd47e10a..86347f44 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -169,7 +169,7 @@ WH_SETTINGS_DIR=/etc/wavehouse/settings WaveHouse keeps all embedded state under a single configurable root, `WH_DATA_DIR` (yaml: `data_dir`). Subdirectories are convention, not config: - `/nats` — embedded NATS JetStream. Holds in-flight events between an ingest POST and the ingest worker → ClickHouse flush, plus the `stream.gap_window_minutes` window (settings directory) of history that powers SSE gap-fill across restarts. -- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). +- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). It grows with every id kept: with `dedupe.retention` at `"0"` (forever) nothing is ever removed, so size the volume for it or set a [retention](/settings-directory#deduplication), whose expired ids an hourly sweep deletes. In a Docker / Podman / Kubernetes deployment, **`data_dir` must resolve to a host-backed volume**. The reference compose file `deployments/compose/standalone.yaml` sets `WH_DATA_DIR=/app/data` and binds a `wavehouse-data:/app/data` volume — copy that pattern. The bundled Dockerfiles pre-create `/app/data` and `/app/settings` owned by the nonroot user (UID 65532); the binary creates the `nats/` and `pebble/` subdirectories under `/app/data` itself on first run. @@ -419,7 +419,9 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres ## Upgrading across the dedupe key change -The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread; nothing removes them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep that will). Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys are never read, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. + +The same release adds **`dedupe.retention`, a required key**: every `config.json`, each tenant's folder included, must state it or the directory is refused (at boot) or not adopted (on reload). `"retention": "0"` keeps every id forever, as before; see [Deduplication](/settings-directory#deduplication) for a finite one. ## Upgrading across the v2 ingest envelope diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 5a736b67..a0a62952 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,9 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. +With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads and deletes 1,024 keys at a time, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. + +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. ## Check your storage before you trust it diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 1e898fdc..126f35dd 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -123,7 +123,8 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `dedupe.enabled` | `false` | Turn deduplication on; a reload opens or closes this tenant's store — see [Deduplication](#deduplication). | | `dedupe.id_field` | `event_id` | Dedup key field — see [Deduplication](#deduplication). | | `dedupe.require_id` | `false` | Reject rows missing the id field — see [Deduplication](#deduplication). | -| `dedupe.tables.
.{id_field, require_id}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | +| `dedupe.retention` | `"0"` | How long a committed id stays a duplicate, as a duration (`"720h"`); `"0"` keeps it forever — see [Deduplication](#deduplication). | +| `dedupe.tables.
.{id_field, require_id, retention}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | | `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the tenant's dead-letter stream (`DLQ_{tenant}`) (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | | `dlq.tables.
.enabled` | `{}` | Optional per-table override of the switch. | | `query.timestamp_bucket_seconds` | `60` | Bucket (seconds, `>= 0`) that a structured query's relative time range is truncated to, so near-identical queries share a cache entry; `0` disables bucketing. Read per query. | @@ -161,8 +162,9 @@ The tenant tunables. Every key is required (a missing one is a validation error) "enabled": false, "id_field": "event_id", "require_id": false, + "retention": "720h", "tables": { - "clicks": { "id_field": "click_id" } + "clicks": { "id_field": "click_id", "retention": "24h" } } }, "dlq": { @@ -188,7 +190,8 @@ Every dedupe knob lives here — there are no boot-config keys for it. The switc - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. +- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep deletes the expired id from the store: first a minute after the store opens, then hourly, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"` or a bare number. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 6f129e34..1189547a 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -50,12 +50,11 @@ type IngestHandler struct { // store, picked off the store the handler already holds (#583 story 7; // dedupe.Stores in production). nil when no dedupe store is wired (tests). Dedup func(store *settings.Store) dedupe.Deduplicator - // DedupeSettings resolves the effective dedupe id_field/require_id for a - // table of the request's tenant ((*settings.Store).DedupeFor in - // production). Called once per record so a settings reload lands at a - // record boundary — one record never mixes two documents' values. Dedup is - // skipped when nil. - DedupeSettings func(store *settings.Store, table string) (enabled bool, idField string, requireID bool) + // DedupeSettings resolves the effective dedupe settings for a table of the + // request's tenant ((*settings.Store).DedupeFor in production). Called + // once per record so a settings reload lands at a record boundary — one + // record never mixes two documents' values. Dedup is skipped when nil. + DedupeSettings func(store *settings.Store, table string) settings.Dedupe // DedupeLease is how long a record's claimed id stays pending while it is // published; 0 means dedupe.DefaultLease. DedupeLease time.Duration @@ -572,8 +571,10 @@ type pendingRecord struct { reject *recordReject // non-nil: the record is bad and is not published payload []byte // the encoded envelope to publish // key is the record's dedupe identity, nil when it is published - // un-deduped; claim is Reserve's answer for it. + // un-deduped; retention is how long its id stays a duplicate once + // committed; claim is Reserve's answer for it. key *dedupe.Key + retention time.Duration claim dedupe.Claim duplicate bool } @@ -701,7 +702,7 @@ func (h *IngestHandler) prepareRecord( // enforces) after the permission checks: check clauses keep pre-#372 semantics. h.validator().CanonicalizeTimestamps(schema, data) - // Optional deduplication. enabled/id_field/require_id resolve per record + // Optional deduplication. The dedupe settings resolve per record // from one snapshot (table override → global; the settings directory // always states them, so no compiled fallback is needed), so a reload // lands at a record boundary. A Deduplicator without a settings source is @@ -709,14 +710,16 @@ func (h *IngestHandler) prepareRecord( // in ingestWindow, once every record of the window is encoded, so nothing // but the publish can fail while the claim is held. if h.Dedup != nil && h.DedupeSettings != nil { - if enabled, idField, requireID := h.DedupeSettings(store, table); enabled { + if dd := h.DedupeSettings(store, table); dd.Enabled { + idField := dd.IDField // An explicit null is as missing as an absent key (#370): fmt.Sprint // would make every null "", one id for every such record. if idVal, ok := data[idField]; ok && idVal != nil { rec.key = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} + rec.retention = dd.Retention } else { dedupeMissingIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) - if requireID { + if dd.RequireID { slog.WarnContext(ctx, "dedupe id_field missing or null; rejecting", "id_field", idField, "table", table) return pendingRecord{reject: &recordReject{ Status: http.StatusBadRequest, @@ -793,7 +796,7 @@ func (h *IngestHandler) ingestWindow(ctx context.Context, store *settings.Store, return h.publishFailed(ctx, dd, topic, recs, i, err) } } - commitClaims(ctx, dd, claimedIn(recs), table) + commitClaims(ctx, dd, recs, table) return nil } @@ -865,7 +868,7 @@ func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, tab // the first copy landed. The records after k were never sent and are released. func (h *IngestHandler) publishFailed(ctx context.Context, dd dedupe.Deduplicator, topic mq.Topic, recs []pendingRecord, k int, err error) *requestAbort { definite := errors.Is(err, mq.ErrQueueFull) - commitClaims(ctx, dd, claimedIn(recs[:k]), topic.Table) + commitClaims(ctx, dd, recs[:k], topic.Table) after := k + 1 if definite { after = k @@ -890,20 +893,33 @@ func claimedIn(recs []pendingRecord) []dedupe.Claim { return out } -// commitClaims makes published records' ids duplicates. A failure does not -// fail the records — they are in the queue — so it is logged and counted, and -// the claims lapse after their lease. -func commitClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim, table string) { - if len(claims) == 0 { - return +// commitClaims makes the ids of recs' Claimed claims duplicates, one Commit +// per retention — one in practice, unless a reload changed it mid-window. A +// failure does not fail the records — they are in the queue — so it is logged +// and counted, and the claims lapse after their lease. +func commitClaims(ctx context.Context, dd dedupe.Deduplicator, recs []pendingRecord, table string) { + var retentions []time.Duration + byRetention := map[time.Duration][]dedupe.Claim{} + for i := range recs { + if recs[i].claim.Status != dedupe.Claimed { + continue + } + r := recs[i].retention + if _, ok := byRetention[r]; !ok { + retentions = append(retentions, r) + } + byRetention[r] = append(byRetention[r], recs[i].claim) } - // The records are queued whatever the request's context does next. - err := dd.Commit(context.WithoutCancel(ctx), claims, 0) - switch { - case err == nil, errors.Is(err, dedupe.ErrDisabled): - default: - dedupeCommitFailedCounter.Add(ctx, int64(len(claims)), metric.WithAttributes(attribute.String("table", table))) - slog.ErrorContext(ctx, "dedupe commit failed after publish; the ids lapse with their lease", "error", err, "table", table, "records", len(claims)) + for _, r := range retentions { + claims := byRetention[r] + // The records are queued whatever the request's context does next. + err := dd.Commit(context.WithoutCancel(ctx), claims, r) + switch { + case err == nil, errors.Is(err, dedupe.ErrDisabled): + default: + dedupeCommitFailedCounter.Add(ctx, int64(len(claims)), metric.WithAttributes(attribute.String("table", table))) + slog.ErrorContext(ctx, "dedupe commit failed after publish; the ids lapse with their lease", "error", err, "table", table, "records", len(claims)) + } } } diff --git a/internal/api/ingest_retention_test.go b/internal/api/ingest_retention_test.go new file mode 100644 index 00000000..ef65e4b2 --- /dev/null +++ b/internal/api/ingest_retention_test.go @@ -0,0 +1,104 @@ +package api + +import ( + "net/http" + "net/http/httptest" + "os" + "path/filepath" + "strings" + "sync/atomic" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/discovery" + "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil" +) + +// A finite retention must outlast the queue's duplicate window, or an id +// re-sent after it expires is claimed again and then dropped by the queue as +// a copy of the first publish. +func TestIngest_MinDedupeRetentionCoversTheDuplicateWindow(t *testing.T) { + t.Parallel() + assert.GreaterOrEqual(t, settings.MinDedupeRetention, mq.EmbeddedDuplicateWindow) +} + +// dedupeConfig is fullConfig with dedupe switched on and the given dedupe +// block's retention settings. +func dedupeConfig(retention, tables string) string { + return strings.Replace(fullConfig(100), + `"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}`, + `"dedupe": {"enabled": true, "id_field": "event_id", "require_id": false, "retention": "`+retention+`", "tables": `+tables+`}`, 1) +} + +// Each record is committed with its table's retention from the adopted +// settings, and a reload changes it for the next request: the retention is +// read per record, like id_field, not fixed when the store was opened. +func TestIngest_Dedup_CommitsWithTheAdoptedRetention(t *testing.T) { + t.Parallel() + dir := writeSettingsFixture(t, dedupeConfig("720h", `{"users": {"retention": "0"}}`)) + tenants, findings := settings.Open(dir) + require.NotNil(t, tenants, "findings: %v", findings) + store, _ := tenants.For(tenant.Default) + + reg := testutil.NewTestSchemaRegistry(t, []*discovery.TableSchema{ + {Name: "clicks", Columns: []discovery.Column{{Name: "event_id", Type: "String"}}}, + {Name: "users", Columns: []discovery.Column{{Name: "event_id", Type: "String"}}}, + }) + dedup := testutil.NewMockDeduplicator() + h := NewIngestHandler(fixedRegistry(reg), &testutil.MockPublisher{}) + h.Dedup = staticDedup(dedup) + h.DedupeSettings = (*settings.Store).DedupeFor + ingest := func(table, id string) { + t.Helper() + w := httptest.NewRecorder() + req := ingestRequest(t, table, map[string]any{"event_id": id}) + h.Handle(w, req.WithContext(WithStore(req.Context(), store))) + require.Equal(t, http.StatusOK, w.Code, w.Body.String()) + } + + ingest("clicks", "e1") + ingest("users", "e1") + assert.Equal(t, 720*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "e1"})) + assert.Equal(t, time.Duration(0), dedup.Retention(dedupe.Key{Table: "users", ID: "e1"}), "the table keeps ids forever") + + require.NoError(t, os.WriteFile(filepath.Join(dir, settings.FileConfig), []byte(dedupeConfig("24h", `{}`)), 0o600)) + _, adopted := tenants.Reload("test") + require.True(t, adopted) + ingest("clicks", "e2") + ingest("users", "e2") + assert.Equal(t, 24*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "e2"})) + assert.Equal(t, 24*time.Hour, dedup.Retention(dedupe.Key{Table: "users", ID: "e2"}), "the override is gone") + assert.Equal(t, 720*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "e1"}), "ids committed before the change keep theirs") +} + +// A reload that lands mid-window splits the window's commit by retention, so +// every record keeps the retention of the snapshot it was prepared under. +func TestIngest_Dedup_ReloadMidWindowCommitsEachRetention(t *testing.T) { + t.Parallel() + dedup := testutil.NewMockDeduplicator() + h := NewIngestHandler(fixedRegistry(testRegistry(t)), &testutil.MockPublisher{}) + h.Dedup = staticDedup(dedup) + var calls atomic.Int32 + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + if calls.Add(1) <= 2 { + return settings.Dedupe{Enabled: true, IDField: "event_id", Retention: time.Hour} + } + return settings.Dedupe{Enabled: true, IDField: "event_id", Retention: 2 * time.Hour} + } + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", + `{"page": "/", "event_id": "a"}`, `{"page": "/", "event_id": "b"}`, `{"page": "/", "event_id": "c"}`))) + require.Equal(t, http.StatusOK, w.Code, w.Body.String()) + assert.Equal(t, 2, dedup.Commits, "one Commit per retention") + assert.Equal(t, time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "a"})) + assert.Equal(t, time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "b"})) + assert.Equal(t, 2*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "c"})) +} diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index 57993333..763ba283 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -201,7 +201,9 @@ func TestIngest_Dedup_FirstTime(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "evt-1"}) w := httptest.NewRecorder() @@ -217,7 +219,9 @@ func TestIngest_Dedup_Duplicate(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } // First call. req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "dup-1"}) @@ -703,7 +707,9 @@ func TestIngest_DedupIsTheTenants(t *testing.T) { pub := &testutil.MockPublisher{} h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = func(s *settings.Store) dedupe.Deduplicator { return stores.For(s.Tenant()) } - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } ingest := func(id tenant.ID) string { store, ok := tenants.For(id) @@ -726,7 +732,9 @@ func TestIngest_Dedup_MissingIDField(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } // Payload omits event_id and require_id is off: the row skips // dedup and is still published — the warn+counter path, not a rejection (#219). @@ -745,7 +753,9 @@ func TestIngest_Dedup_RequireID_Rejects(t *testing.T) { pub := &testutil.MockPublisher{} h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(testutil.NewMockDeduplicator()) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: true} + } w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home"}))) @@ -767,7 +777,9 @@ func TestIngest_NDJSON_RequireID_Rejects(t *testing.T) { pub := &testutil.MockPublisher{} h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(testutil.NewMockDeduplicator()) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: true} + } req := ndjsonRequest(t, "clicks", jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), @@ -1015,7 +1027,9 @@ func TestIngest_NDJSON_Dedup(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } req := ndjsonRequest(t, "clicks", jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), @@ -2306,7 +2320,9 @@ func TestIngest_Dedup_DisabledBySettings(t *testing.T) { dedup.Err = errors.New("must not be called while disabled") h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return false, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{IDField: "event_id", RequireID: true} + } w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", tt.body))) @@ -2326,7 +2342,9 @@ func TestIngest_Dedup_DisabledMidReload(t *testing.T) { dedup.Err = dedupe.ErrDisabled h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: true} + } w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"event_id": "e1", "page": "/home"}))) @@ -2744,7 +2762,9 @@ func dedupHandler(t *testing.T, pub *testutil.MockPublisher, dedup dedupe.Dedupl t.Helper() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", requireID } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: requireID} + } return h } diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index 3d1dc931..32015dc4 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -391,7 +391,9 @@ func pebbleBatchHandler(tb testing.TB, window int) (*IngestHandler, *countingDed counted := &countingDedup{Deduplicator: store} h := NewIngestHandler(fixedRegistry(testRegistry(tb)), &testutil.MockPublisher{}) h.Dedup = staticDedup(counted) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } h.window = window return h, counted } diff --git a/internal/api/settings_test.go b/internal/api/settings_test.go index 9e823cea..24fa80f1 100644 --- a/internal/api/settings_test.go +++ b/internal/api/settings_test.go @@ -19,7 +19,7 @@ import ( // fullConfig is a complete config.json (every key is required) with the // given query.default_max_rows. func fullConfig(maxRows int) string { - return fmt.Sprintf(`{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": %d, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, maxRows) + return fmt.Sprintf(`{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": %d, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, maxRows) } // writeSettingsFixture materializes a minimal valid settings directory whose diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 1468d0df..019214c0 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -208,7 +208,7 @@ func TestNew_DedupeFollowsSettings(t *testing.T) { for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { dir := writeSettings(t, map[string]any{"dedupe": map[string]any{ - "enabled": tt.enabled, "id_field": "event_id", "require_id": false, "tables": map[string]any{}, + "enabled": tt.enabled, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}, }}) cfg := testConfig(t, dir) a := newApp(t, cfg, Options{}) @@ -241,7 +241,7 @@ func TestReload_DrivesTheRegisteredHooks(t *testing.T) { require.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, dir, map[string]any{ - "dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}, + "dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}, "mq": map[string]any{"max_bytes_gb": 2}, }) _, adopted := a.tenants.Reload("test") @@ -405,7 +405,7 @@ func TestNew_NestedWithoutAnOperatorKeyWarnsTheOpsTreeIsClosed(t *testing.T) { // request, so a lost 0 folder is felt at once on the routes that read tenant // 0's list. func TestReload_NestedHooksFollowEachTenant(t *testing.T) { - dedupeOn := map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}} + dedupeOn := map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}} grown := map[string]any{"dedupe": dedupeOn, "mq": map[string]any{"max_bytes_gb": 2}} root := writeNestedSettings(t, map[string]map[string]any{ "0": {"mq": map[string]any{"max_bytes_gb": 1}}, @@ -478,7 +478,7 @@ func TestReload_NestedHooksFollowEachTenant(t *testing.T) { // reopened over the same seen ids when the folder is back. The instance is // open while some tenant's store is, and Close releases it. func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { - dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}} + dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}} root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil, "broken": invalidQuery}) cfg := testConfig(t, root) a := newApp(t, cfg, Options{}) @@ -542,7 +542,7 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { // instance — their ingest answers 500 until a reload or a restart opens it — // while the process, and every tenant with dedupe off, carries on. func TestNew_DedupeOpenFailure(t *testing.T) { - dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}} + dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}} // A regular file where the instance's directory should be is what Pebble // refuses to open. block := func(t *testing.T, dataDir string) { @@ -852,7 +852,7 @@ func analystPipe(t *testing.T, dir string) { func TestNew_LateBootFailureReleasesEverything(t *testing.T) { guardGlobals(t) dir := writeSettings(t, map[string]any{"dedupe": map[string]any{ - "enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}, + "enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}, }}) cfg := testConfig(t, dir) natsDir := filepath.Join(cfg.DataDir, "nats") @@ -1458,7 +1458,7 @@ func TestReload_CeilingRefusesAThirdTupleThenOpensIt(t *testing.T) { func TestReload_TenantGoneReleasesItsPoolAndRegistry(t *testing.T) { jwks, _, fetches := jwksServer(t, "acme-1") acmeSettings := authPatch(jwks.URL) - acmeSettings["dedupe"] = map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}} + acmeSettings["dedupe"] = map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}} root := writeNestedSettings(t, map[string]map[string]any{"acme": acmeSettings, "globex": nil}) a := newApp(t, testConfig(t, root), Options{}) acme, acmeRegistry, acmeDedup := a.pools.For("acme"), a.discoveries.For("acme"), a.dedup.For("acme") diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index b3f015f0..5186a738 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -30,9 +30,16 @@ import ( type Embedded struct { dir string - mu sync.Mutex // guards db and open - db *pebble.DB - open int // tenant stores open over db + mu sync.Mutex // guards db, open and stopSweep + db *pebble.DB + open int // tenant stores open over db + stopSweep func() // stops db's sweep + + // commitMu is read-held by Commit and held by a sweep chunk, so a sweep + // never deletes a key a Commit rewrote after the sweep read it. + commitMu sync.RWMutex + sweepFirst time.Duration + sweepEvery time.Duration pending *pendingSet tokens atomic.Uint64 @@ -45,7 +52,13 @@ type Embedded struct { // NewEmbedded returns the embedded implementation under dataDir. Nothing is // opened until a tenant's store is. func NewEmbedded(dataDir string) *Embedded { - return &Embedded{dir: filepath.Join(dataDir, "pebble"), pending: newPendingSet(), now: time.Now} + return &Embedded{ + dir: filepath.Join(dataDir, "pebble"), + pending: newPendingSet(), + now: time.Now, + sweepFirst: sweepFirstDelay, + sweepEvery: sweepInterval, + } } // Dir is where the instance lives. @@ -76,6 +89,7 @@ func (e *Embedded) acquire(prefix []byte) (Deduplicator, error) { return nil, err } e.db = db + e.stopSweep = e.startSweep(db) } e.open++ return &tenantStore{e: e, db: e.db, prefix: prefix}, nil @@ -90,6 +104,7 @@ func (e *Embedded) release() error { if e.open > 0 { return nil } + e.stopSweep() err := e.db.Close() e.db = nil return err @@ -182,24 +197,36 @@ func (s *tenantStore) reserve(key []byte, k Key, now time.Time, lease time.Durat // committedLive reports whether a stored value is a commit that has not // expired. func committedLive(val []byte, now time.Time) bool { + exp, ok := committedExpiry(val) + return ok && (exp == 0 || now.UnixNano() < exp) +} + +// committedExpired reports whether a stored value is a commit whose +// retention has ended — what the sweep deletes. +func committedExpired(val []byte, now time.Time) bool { + exp, ok := committedExpiry(val) + return ok && exp != 0 && now.UnixNano() >= exp +} + +// committedExpiry reads a commit's expiry (UnixNano, 0 = never); ok is false +// for a value that is not a commit. +func committedExpiry(val []byte) (exp int64, ok bool) { if len(val) != valueLen || val[0] != committedMark { - return false + return 0, false } - exp := int64(binary.BigEndian.Uint64(val[1:])) //nolint:gosec // written from an int64 below - return exp == 0 || now.UnixNano() < exp + return int64(binary.BigEndian.Uint64(val[1:])), true //nolint:gosec // written from an int64 below } // Commit writes every claim in one batch and one fsync, then drops the // pending entries it still owns — in that order, so no Reserve in between // finds the key neither pending nor committed. func (s *tenantStore) Commit(_ context.Context, claims []Claim, retention time.Duration) error { - var exp int64 - if retention > 0 { - exp = s.e.now().Add(retention).UnixNano() - } + s.e.commitMu.RLock() + defer s.e.commitMu.RUnlock() + exp := expiry(s.e.now(), retention) val := make([]byte, valueLen) val[0] = committedMark - binary.BigEndian.PutUint64(val[1:], uint64(exp)) + binary.BigEndian.PutUint64(val[1:], uint64(exp)) //nolint:gosec // expiry is never negative b := s.db.NewBatch() defer func() { _ = b.Close() }() for _, c := range claims { @@ -214,6 +241,20 @@ func (s *tenantStore) Commit(_ context.Context, claims []Claim, retention time.D return nil } +// expiry is the stored expiry of a commit at now kept for retention: 0 for +// none, and the latest representable instant for a retention reaching past +// it, rather than a wrapped-around one in the past. +func expiry(now time.Time, retention time.Duration) int64 { + if retention <= 0 { + return 0 + } + n := now.UnixNano() + if retention > time.Duration(math.MaxInt64-n) { + return math.MaxInt64 + } + return n + int64(retention) +} + // Release drops the pending entries the claims still own. func (s *tenantStore) Release(_ context.Context, claims []Claim) error { s.release(claims) diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go new file mode 100644 index 00000000..e536c5ed --- /dev/null +++ b/internal/dedupe/sweep.go @@ -0,0 +1,151 @@ +package dedupe + +import ( + "bytes" + "context" + "fmt" + "log/slog" + "time" + + "github.com/cockroachdb/pebble" + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/metric" +) + +// The sweep's cadence. Expired keys are already absent to Reserve, so the +// sweep only reclaims space and can run rarely; the first pass comes soon +// after the instance opens so an upgrade's version-0 keys go without waiting +// an hour. +const ( + sweepInterval = time.Hour + sweepFirstDelay = time.Minute + // sweepChunk keys are read and deleted per lock hold, with sweepPause + // between chunks: at most ~100k keys a second, and a Commit never waits + // longer than one chunk. + sweepChunk = 1024 + sweepPause = 10 * time.Millisecond +) + +// Swept-key reasons, the metric's reason attribute. +const ( + sweptExpired = "expired" + sweptVersion0 = "version_0" + sweptAttribute = "reason" +) + +var sweptKeysCounter, _ = otel.Meter("wavehouse-dedupe").Int64Counter( + "wavehouse_dedupe_swept_keys_total", + metric.WithDescription("Keys the embedded dedupe sweep deleted, by reason: expired (retention ended) or version_0 (the layout before ids were keyed by table)"), +) + +// sweepResult is what a sweep deleted. +type sweepResult struct { + Expired, Version0 int +} + +// startSweep runs the sweep over db until the returned stop is called; stop +// waits for a chunk in progress to finish. Callers hold e.mu. +func (e *Embedded) startSweep(db *pebble.DB) (stop func()) { + ctx, cancel := context.WithCancel(context.Background()) + done := make(chan struct{}) + go func() { + defer close(done) + wait := e.sweepFirst + for { + select { + case <-ctx.Done(): + return + case <-time.After(wait): + } + wait = e.sweepEvery + res, err := e.sweep(ctx, db) + switch { + case err != nil && ctx.Err() == nil: + slog.WarnContext(ctx, "dedupe sweep failed; retrying next interval", "error", err, "expired", res.Expired, "version_0", res.Version0) + case res.Expired+res.Version0 > 0: + slog.InfoContext(ctx, "dedupe sweep deleted keys", "expired", res.Expired, "version_0", res.Version0) + } + } + }() + return func() { + cancel() + <-done + } +} + +// sweep makes one pass over the whole instance, deleting keys whose +// retention has ended and version-0 keys (tenant ‖ 0x00 ‖ id, from before +// ids were keyed by table), which nothing reads. It stops early, without +// error, when ctx ends. +func (e *Embedded) sweep(ctx context.Context, db *pebble.DB) (sweepResult, error) { + var res sweepResult + var from []byte + for { + next, err := e.sweepChunk(ctx, db, from, &res) + if err != nil || next == nil { + return res, err + } + from = next + select { + case <-ctx.Done(): + return res, nil + case <-time.After(sweepPause): + } + } +} + +// sweepChunk deletes the sweepable keys among the next sweepChunk keys from +// from, returning where the next chunk starts (nil at the end). It holds +// commitMu, so no Commit lands between reading a key and deleting it: a key +// re-committed after it expired is never deleted with its new value. +func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, res *sweepResult) ([]byte, error) { + e.commitMu.Lock() + defer e.commitMu.Unlock() + now := e.now() + it, err := db.NewIter(&pebble.IterOptions{LowerBound: from}) + if err != nil { + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + b := db.NewBatch() + defer func() { _ = b.Close() }() + var next []byte + var expired, v0 int64 + seen := 0 + for valid := it.First(); valid; valid = it.Next() { + if seen == sweepChunk { + next = bytes.Clone(it.Key()) + break + } + seen++ + k := it.Key() + switch { + case len(k) == 0 || k[0] != keyVersion: + v0++ + case committedExpired(it.Value(), now): + expired++ + default: + continue + } + if err := b.Delete(k, nil); err != nil { + _ = it.Close() + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + } + if err := it.Close(); err != nil { + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + // NoSync: a delete lost to a crash is redone by the next pass. + if err := b.Commit(pebble.NoSync); err != nil { + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + res.Expired += int(expired) + res.Version0 += int(v0) + if expired > 0 { + sweptKeysCounter.Add(ctx, expired, metric.WithAttributes(attribute.String(sweptAttribute, sweptExpired))) + } + if v0 > 0 { + sweptKeysCounter.Add(ctx, v0, metric.WithAttributes(attribute.String(sweptAttribute, sweptVersion0))) + } + return next, nil +} diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go new file mode 100644 index 00000000..44579dcd --- /dev/null +++ b/internal/dedupe/sweep_test.go @@ -0,0 +1,158 @@ +package dedupe + +import ( + "context" + "errors" + "fmt" + "math" + "sync/atomic" + "testing" + "time" + + "github.com/cockroachdb/pebble" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// stepClock is a clock a test moves by hand. +type stepClock struct{ ns atomic.Int64 } + +func newStepClock() *stepClock { + c := &stepClock{} + c.ns.Store(time.Now().UnixNano()) + return c +} + +func (c *stepClock) now() time.Time { return time.Unix(0, c.ns.Load()) } +func (c *stepClock) advance(d time.Duration) { c.ns.Add(int64(d)) } +func present(t *testing.T, e *Embedded, key []byte) bool { + t.Helper() + _, closer, err := e.db.Get(key) + if errors.Is(err, pebble.ErrNotFound) { + return false + } + require.NoError(t, err) + _ = closer.Close() + return true +} + +// commitIDs reserves and commits ids in table "events" with retention. +func commitIDs(t *testing.T, m *Managed, retention time.Duration, ids ...string) { + t.Helper() + keys := make([]Key, len(ids)) + for i, id := range ids { + keys[i] = Key{Table: "events", ID: id} + } + claims, err := m.Reserve(context.Background(), keys, DefaultLease) + require.NoError(t, err) + require.NoError(t, m.Commit(context.Background(), claims, retention)) +} + +// A sweep deletes the keys whose retention has ended and every version-0 +// key, across chunk boundaries, and leaves every live key: one kept forever, +// one not yet expired, and one that expired and was committed again. +func TestEmbedded_SweepDeletesExpiredAndVersionZeroKeys(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + clock := newStepClock() + SetClock(e, clock.now) + acme, globex := switchedOn(t, e, "acme"), switchedOn(t, e, "globex") + + // More expired keys than one chunk holds, interleaved with live ones. + var expired, live []string + for i := range 2*sweepChunk + 10 { + expired = append(expired, fmt.Sprintf("x%05d", i)) + live = append(live, fmt.Sprintf("x%05d-live", i)) + } + commitIDs(t, acme, time.Hour, expired...) + commitIDs(t, acme, 3*time.Hour, live...) + commitIDs(t, globex, 0, "forever") + commitIDs(t, globex, time.Hour, "recommitted") + for _, k := range []string{"acme\x00e1", "acme\x00e2", "globex\x00e1"} { + require.NoError(t, e.db.Set([]byte(k), make([]byte, 8), pebble.Sync)) + } + + clock.advance(2 * time.Hour) + commitIDs(t, globex, time.Hour, "recommitted") + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Expired: len(expired), Version0: 3}, res) + + for _, id := range expired { + require.False(t, present(t, e, AppendKey(nil, KeyPrefix("acme"), Key{Table: "events", ID: id})), id) + } + for _, id := range live { + require.True(t, present(t, e, AppendKey(nil, KeyPrefix("acme"), Key{Table: "events", ID: id})), id) + } + assert.True(t, present(t, e, AppendKey(nil, KeyPrefix("globex"), Key{Table: "events", ID: "forever"}))) + assert.True(t, present(t, e, AppendKey(nil, KeyPrefix("globex"), Key{Table: "events", ID: "recommitted"}))) + assert.False(t, present(t, e, []byte("acme\x00e1"))) + assert.False(t, present(t, e, []byte("globex\x00e1"))) + + dup, err := mark(context.Background(), globex, "recommitted") + require.NoError(t, err) + assert.True(t, dup, "the new commit survived the sweep") + res, err = e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{}, res, "a second pass finds nothing") +} + +// A retention is honoured on read before any sweep has run: the key is a +// duplicate until the retention ends and claimable from that instant. +func TestEmbedded_RetentionHonouredOnRead(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + clock := newStepClock() + SetClock(e, clock.now) + m := switchedOn(t, e, "acme") + commitIDs(t, m, time.Hour, "e1") + + clock.advance(time.Hour - time.Nanosecond) + claims, err := m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) + require.NoError(t, err) + assert.Equal(t, Duplicate, claims[0].Status) + + clock.advance(time.Nanosecond) + claims, err = m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) + require.NoError(t, err) + assert.Equal(t, Claimed, claims[0].Status) +} + +// The sweep runs on its own once the instance opens, and stops with it. +func TestEmbedded_SweepRunsWhileOpen(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + e.sweepFirst, e.sweepEvery = time.Millisecond, time.Millisecond + m := e.Tenant("acme") + require.NoError(t, m.Apply(true)) + require.NoError(t, e.db.Set([]byte("acme\x00e1"), make([]byte, 8), pebble.Sync)) + assert.Eventually(t, func() bool { return !present(t, e, []byte("acme\x00e1")) }, 5*time.Second, 5*time.Millisecond) + require.NoError(t, m.Apply(false), "closing waits for the sweep to stop") + assert.False(t, e.Open()) +} + +// A sweep stops between chunks when its context ends. +func TestEmbedded_SweepStopsWhenCancelled(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + switchedOn(t, e, "acme") + b := e.db.NewBatch() + for i := range 3 * sweepChunk { + require.NoError(t, b.Set(fmt.Appendf(nil, "acme\x00%05d", i), nil, nil)) + } + require.NoError(t, b.Commit(pebble.Sync)) + ctx, cancel := context.WithCancel(context.Background()) + cancel() + res, err := e.sweep(ctx, e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Version0: sweepChunk}, res, "one chunk, then the cancellation is seen") +} + +func TestExpiry(t *testing.T) { + t.Parallel() + now := time.Unix(0, 1_000) + assert.Zero(t, expiry(now, 0)) + assert.Zero(t, expiry(now, -time.Second)) + assert.Equal(t, 1_000+int64(time.Hour), expiry(now, time.Hour)) + assert.Equal(t, int64(math.MaxInt64), expiry(now, time.Duration(math.MaxInt64)), "saturates rather than wrapping into the past") +} diff --git a/internal/settings/registry_test.go b/internal/settings/registry_test.go index 06e2a3b9..3ac222cc 100644 --- a/internal/settings/registry_test.go +++ b/internal/settings/registry_test.go @@ -113,9 +113,7 @@ func TestRegistry_SurvivesVanishedDirectory(t *testing.T) { assert.False(t, adopted) assert.True(t, HasErrors(findings)) assert.Equal(t, 42, s.DefaultMaxRows()) - _, id, req := s.DedupeFor("clicks") - assert.Equal(t, "event_id", id) - assert.False(t, req) + assert.Equal(t, "event_id", s.DedupeFor("clicks").IDField) } // TestRegistry_AfterAdoptRunsOnlyOnAdoption pins the lifecycle hook contract diff --git a/internal/settings/seed/config.json b/internal/settings/seed/config.json index a8ab41a2..61aa8fee 100644 --- a/internal/settings/seed/config.json +++ b/internal/settings/seed/config.json @@ -26,6 +26,7 @@ "enabled": false, "id_field": "event_id", "require_id": false, + "retention": "0", "tables": {} }, "dlq": { diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 7da6c3fa..68445543 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -13,6 +13,8 @@ package settings import ( + "time" + "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" ) @@ -141,14 +143,18 @@ type AuthConfig struct { // (dedupe.Managed, one per tenant, each a share of the one embedded Pebble // instance), so the whole block is tenant-owned. // -// id_field and require_id are required here and optional per table: a table -// override inherits whichever field it doesn't name. An empty, +// id_field, require_id and retention are required here and optional per +// table: a table override inherits whichever field it doesn't name. An empty, // whitespace-only, or whitespace-padded id_field is rejected at every level, // so the effective id_field can never be empty or silently unmatchable. type DedupeConfig struct { Enabled *bool `json:"enabled"` IDField *string `json:"id_field"` RequireID *bool `json:"require_id"` + // Retention is how long a committed id stays a duplicate, as a Go + // duration ("720h"); "0" keeps it forever. A change applies to ids + // committed after it. + Retention *string `json:"retention"` // Tables holds per-table overrides keyed by ClickHouse table name (#222). // Names are format-checked only — existence is schema discovery's runtime // concern, same as policies.json table keys. @@ -160,8 +166,16 @@ type DedupeConfig struct { type TableDedupe struct { IDField *string `json:"id_field,omitempty"` RequireID *bool `json:"require_id,omitempty"` + Retention *string `json:"retention,omitempty"` } +// MinDedupeRetention is the shortest finite dedupe retention: the embedded +// queue's duplicate window (mq.EmbeddedDuplicateWindow). A record is +// published under an idempotency key derived from its id, so an id re-sent +// after a shorter retention but inside the window is claimed again and then +// dropped by the queue as a copy, while the client is told it was accepted. +const MinDedupeRetention = 2 * time.Minute + // DLQConfig gates the Dead Letter Queue: whether a row that still fails // after the row-by-row isolation retry is parked on the tenant's dead-letter // queue (and its original acked) or left unacked to be redelivered diff --git a/internal/settings/store.go b/internal/settings/store.go index f68a2fbf..b3615efe 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -76,23 +76,38 @@ func (s *Store) DedupeEnabled() bool { return *s.doc().Config.Dedupe.Enabled } +// Dedupe is a table's effective dedupe settings. +type Dedupe struct { + Enabled bool + IDField string + RequireID bool + // Retention is how long a committed id stays a duplicate; 0 is forever. + Retention time.Duration +} + // DedupeFor resolves the effective dedupe settings for a table: the switch, // then the table override for each field it names, the global value -// otherwise. All three resolve from one snapshot load, so a reload can never -// hand a record the id_field of one document and the require_id (or enabled) -// of another. -func (s *Store) DedupeFor(table string) (enabled bool, idField string, requireID bool) { +// otherwise. Every field resolves from one snapshot load, so a reload can +// never hand a record the id_field of one document and the require_id, +// retention or switch of another. +func (s *Store) DedupeFor(table string) Dedupe { d := s.doc().Config.Dedupe - enabled, idField, requireID = *d.Enabled, *d.IDField, *d.RequireID + out := Dedupe{Enabled: *d.Enabled, IDField: *d.IDField, RequireID: *d.RequireID} + retention := *d.Retention if td, ok := d.Tables[table]; ok { if td.IDField != nil { - idField = *td.IDField + out.IDField = *td.IDField } if td.RequireID != nil { - requireID = *td.RequireID + out.RequireID = *td.RequireID + } + if td.Retention != nil { + retention = *td.Retention } } - return enabled, idField, requireID + // Validate has parsed it already. + out.Retention, _ = time.ParseDuration(retention) + return out } // ClickHouse is the adopted connection wiring, resolved as one value from diff --git a/internal/settings/store_test.go b/internal/settings/store_test.go index abcc6c90..e9689824 100644 --- a/internal/settings/store_test.go +++ b/internal/settings/store_test.go @@ -46,23 +46,22 @@ func TestStore_Tenant(t *testing.T) { func TestStore_DedupeFor_Cascade(t *testing.T) { t.Parallel() s := newLoadedStore(t, map[string]string{ - FileConfig: configJSON(`{"dedupe": {"require_id": true, "tables": {"clicks": {"id_field": "click_id"}, "views": {"require_id": false}}}}`), + FileConfig: configJSON(`{"dedupe": {"require_id": true, "retention": "720h", "tables": {"clicks": {"id_field": "click_id"}, "views": {"require_id": false, "retention": "24h"}, "audit": {"retention": "0"}}}}`), }) tests := []struct { - name, table, wantID string - wantRequire bool + name, table string + want Dedupe }{ - {name: "table overrides id_field, inherits require_id", table: "clicks", wantID: "click_id", wantRequire: true}, - {name: "table overrides require_id, inherits id_field", table: "views", wantID: "event_id", wantRequire: false}, - {name: "unlisted table gets globals", table: "other", wantID: "event_id", wantRequire: true}, + {name: "table overrides id_field, inherits the rest", table: "clicks", want: Dedupe{IDField: "click_id", RequireID: true, Retention: 720 * time.Hour}}, + {name: "table overrides require_id and retention, inherits id_field", table: "views", want: Dedupe{IDField: "event_id", Retention: 24 * time.Hour}}, + {name: "table keeps ids forever under a finite tenant retention", table: "audit", want: Dedupe{IDField: "event_id", RequireID: true}}, + {name: "unlisted table gets globals", table: "other", want: Dedupe{IDField: "event_id", RequireID: true, Retention: 720 * time.Hour}}, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { t.Parallel() - _, id, req := s.DedupeFor(tt.table) - assert.Equal(t, tt.wantID, id) - assert.Equal(t, tt.wantRequire, req) + assert.Equal(t, tt.want, s.DedupeFor(tt.table)) }) } } @@ -83,9 +82,7 @@ func TestStore_SeedIsValid(t *testing.T) { // decision (deployments/compose/settings ships the opt-in trial one). assert.Len(t, findings, 1, "findings: %s", findingStrings(findings)) assert.Contains(t, findingStrings(findings), "no policy") - _, id, req := s.DedupeFor("anything") - assert.Equal(t, "event_id", id) - assert.False(t, req) + assert.Equal(t, Dedupe{IDField: "event_id"}, s.DedupeFor("anything"), "retention 0: ids kept forever, as before retention existed") assert.Equal(t, ClickHouse{Addr: "localhost:9000", HTTPPort: 8123, HTTPScheme: "http", Database: "default", Username: "default", QueryTimeout: 30 * time.Second, Headers: map[string]string{}, MaxOpenConns: 10, MaxIdleConns: 5}, s.ClickHouse()) assert.Equal(t, Auth{JWKSURL: "", RoleClaim: "role"}, s.Auth()) assert.True(t, s.DLQFor("anything")) diff --git a/internal/settings/validate.go b/internal/settings/validate.go index 08575c95..76352e19 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -12,6 +12,7 @@ import ( "path/filepath" "slices" "strings" + "time" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" @@ -450,6 +451,26 @@ func (v *validator) checkIDField(path string, val *string) { } } +// checkRetention rejects a dedupe retention that is not a duration, is +// negative, or is finite but shorter than MinDedupeRetention. The short one +// is refused rather than raised to the minimum, so the file never means +// something other than what it says. nil is the caller's concern, as for +// id_field. +func (v *validator) checkRetention(path string, val *string) { + if val == nil { + return + } + d, err := time.ParseDuration(*val) + switch { + case err != nil: + v.errorf(FileConfig, path, "must be a duration such as \"720h\", or \"0\" to keep ids forever, got %q", *val) + case d < 0: + v.errorf(FileConfig, path, "must not be negative, got %q", *val) + case d > 0 && d < MinDedupeRetention: + v.errorf(FileConfig, path, "%q is shorter than the ingest queue's %s duplicate window: an id re-sent after it expires but inside the window would be dropped by the queue while the client is told it was accepted — use at least %q, or \"0\" to keep ids forever", *val, MinDedupeRetention, MinDedupeRetention.String()) + } +} + // checkTableName rejects a per-table override key that could never match a // table: empty, carrying surrounding whitespace, or holding NUL. Shared by // the dedupe and dlq override maps. @@ -673,15 +694,20 @@ func (v *validator) parseConfig(data []byte) TenantConfig { if d.RequireID == nil { v.required("dedupe.require_id") } + if d.Retention == nil { + v.required("dedupe.retention") + } v.checkIDField("dedupe.id_field", d.IDField) + v.checkRetention("dedupe.retention", d.Retention) // Sorted iteration keeps finding order deterministic across runs. for _, table := range slices.Sorted(maps.Keys(d.Tables)) { td := d.Tables[table] path := "dedupe.tables." + table v.checkTableName("dedupe.tables", table) v.checkIDField(path+".id_field", td.IDField) - if td.IDField == nil && td.RequireID == nil { - v.warnf(FileConfig, path, "override sets nothing — remove it, or set id_field or require_id") + v.checkRetention(path+".retention", td.Retention) + if td.IDField == nil && td.RequireID == nil && td.Retention == nil { + v.warnf(FileConfig, path, "override sets nothing — remove it, or set id_field, require_id or retention") } } } diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index 4dae9c9d..a6d7c5f7 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -269,13 +269,13 @@ func TestValidate_ContentRules(t *testing.T) { {"negative max rows", FileConfig, `{"query": {"default_max_rows": -1}}`, "must be >= 1"}, {"zero max rows", FileConfig, `{"query": {"default_max_rows": 0}}`, "must be >= 1"}, {"missing dedupe block", FileConfig, `{"dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "dedupe: required"}, - {"missing dlq block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "dlq: required"}, + {"missing dlq block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "dlq: required"}, {"missing dlq.enabled", FileConfig, `{"dlq": {"tables": {}}}`, "dlq.enabled: required"}, {"empty dlq override table name", FileConfig, `{"dlq": {"tables": {"": {"enabled": false}}}}`, "table name must not be empty"}, {"dlq override table whitespace", FileConfig, `{"dlq": {"tables": {"clicks ": {"enabled": false}}}}`, "surrounding whitespace"}, {"missing query.timestamp_bucket_seconds", FileConfig, `{"query": {"default_max_rows": 1}}`, "query.timestamp_bucket_seconds: required"}, {"negative timestamp bucket", FileConfig, `{"query": {"timestamp_bucket_seconds": -1}}`, "must be >= 0"}, - {"missing stream block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "cors": {"allowed_origins": []}}`, "stream: required"}, + {"missing stream block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "cors": {"allowed_origins": []}}`, "stream: required"}, {"missing stream.keepalive_interval", FileConfig, `{"stream": {"keepalive_buckets": 3, "gap_window_minutes": 15}}`, "stream.keepalive_interval: required"}, {"missing stream.keepalive_buckets", FileConfig, `{"stream": {"keepalive_interval": 30, "gap_window_minutes": 15}}`, "stream.keepalive_buckets: required"}, {"missing stream.gap_window_minutes", FileConfig, `{"stream": {"keepalive_interval": 30, "keepalive_buckets": 3}}`, "stream.gap_window_minutes: required"}, @@ -284,14 +284,21 @@ func TestValidate_ContentRules(t *testing.T) { {"negative gap window", FileConfig, `{"stream": {"gap_window_minutes": -1}}`, "stream.gap_window_minutes: must be >= 0"}, {"keepalive as a duration string", FileConfig, `{"stream": {"keepalive_interval": "30s"}}`, "keepalive_interval"}, {"missing dedupe.require_id", FileConfig, `{"dedupe": {"id_field": "event_id"}}`, "dedupe.require_id: required"}, - {"missing dedupe.enabled", FileConfig, `{"dedupe": {"id_field": "event_id", "require_id": false}}`, "dedupe.enabled: required"}, + {"missing dedupe.enabled", FileConfig, `{"dedupe": {"id_field": "event_id", "require_id": false, "retention": "0"}}`, "dedupe.enabled: required"}, + {"missing dedupe.retention", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}}`, "dedupe.retention: required"}, + {"dedupe.retention not a duration", FileConfig, configJSON(`{"dedupe": {"retention": "30d"}}`), `dedupe.retention: must be a duration such as "720h"`}, + {"dedupe.retention a number", FileConfig, configJSON(`{"dedupe": {"retention": 3600}}`), "retention"}, + {"dedupe.retention negative", FileConfig, configJSON(`{"dedupe": {"retention": "-1h"}}`), "dedupe.retention: must not be negative"}, + {"dedupe.retention under the duplicate window", FileConfig, configJSON(`{"dedupe": {"retention": "1m59s"}}`), `dedupe.retention: "1m59s" is shorter than the ingest queue's 2m0s duplicate window`}, + {"override retention under the duplicate window", FileConfig, configJSON(`{"dedupe": {"tables": {"clicks": {"retention": "30s"}}}}`), "dedupe.tables.clicks.retention: \"30s\" is shorter"}, + {"override retention not a duration", FileConfig, configJSON(`{"dedupe": {"tables": {"clicks": {"retention": "forever"}}}}`), "dedupe.tables.clicks.retention: must be a duration"}, {"missing query.default_max_rows", FileConfig, `{"query": {}}`, "query.default_max_rows: required"}, {"missing schema.refresh_interval", FileConfig, `{"schema": {}}`, "schema.refresh_interval: required"}, {"missing cors.allowed_origins", FileConfig, `{"cors": {}}`, "cors.allowed_origins: required"}, {"empty config document", FileConfig, `{}`, "cors: required"}, {"otel is boot config", FileConfig, `{"otel": {"enabled": true}}`, "unknown field"}, - {"missing clickhouse block", FileConfig, `{"auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "clickhouse: required"}, - {"missing auth block", FileConfig, `{"clickhouse": {"addr": "h:9000", "http_port": 8123, "http_scheme": "http", "database": "d", "username": "u", "query_timeout": 1}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "auth: required"}, + {"missing clickhouse block", FileConfig, `{"auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "clickhouse: required"}, + {"missing auth block", FileConfig, `{"clickhouse": {"addr": "h:9000", "http_port": 8123, "http_scheme": "http", "database": "d", "username": "u", "query_timeout": 1}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "auth: required"}, {"missing clickhouse.addr", FileConfig, `{"clickhouse": {"http_port": 8123, "http_scheme": "http", "database": "d", "username": "u", "query_timeout": 1}}`, "clickhouse.addr: required"}, {"clickhouse.addr without port", FileConfig, `{"clickhouse": {"addr": "localhost"}}`, "must be host:port"}, {"clickhouse.http_port out of range", FileConfig, `{"clickhouse": {"http_port": 70000}}`, "clickhouse.http_port: must be in 1-65535"}, @@ -344,6 +351,28 @@ func TestValidate_ContentRules(t *testing.T) { } } +// A finite retention at or above the duplicate window is accepted, "0" (or +// any zero duration) is forever, and a table may keep ids longer or shorter +// than the tenant, or forever under a finite tenant retention. +func TestValidate_DedupeRetentionAccepted(t *testing.T) { + t.Parallel() + for _, patch := range []string{ + `{"dedupe": {"retention": "0"}}`, + `{"dedupe": {"retention": "0s"}}`, + `{"dedupe": {"retention": "2m"}}`, + `{"dedupe": {"retention": "720h", "tables": {"clicks": {"retention": "24h"}, "views": {"retention": "0"}}}}`, + } { + t.Run(patch, func(t *testing.T) { + t.Parallel() + files := validFiles() + files[FileConfig] = configJSON(patch) + doc, findings := ValidateDir(writeDir(t, files)) + require.NotNil(t, doc, "findings: %s", findingStrings(findings)) + assert.False(t, HasErrors(findings)) + }) + } +} + // TestValidate_ClickHouseTLSPathsAreNotOpened pins that the tls block is // checked for shape only: Validate is pure and also runs on the control // plane, so paths that exist nowhere still validate, and the values reach diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index 624fe0ef..b82974e7 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -116,6 +116,7 @@ func (m *MockSubscriber) Close() error { return nil } type MockDeduplicator struct { mu sync.Mutex committed map[dedupe.Key]bool + retention map[dedupe.Key]time.Duration // each commit's retention pending map[dedupe.Key]string tokens int // Err, if set, fails Reserve — after ErrAfter calls have succeeded; @@ -132,7 +133,7 @@ type MockDeduplicator struct { var _ dedupe.Deduplicator = (*MockDeduplicator)(nil) func NewMockDeduplicator() *MockDeduplicator { - return &MockDeduplicator{committed: map[dedupe.Key]bool{}, pending: map[dedupe.Key]string{}} + return &MockDeduplicator{committed: map[dedupe.Key]bool{}, retention: map[dedupe.Key]time.Duration{}, pending: map[dedupe.Key]string{}} } // Reserve answers Duplicate for a key repeated in one call, as Managed does. @@ -163,7 +164,7 @@ func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time. return claims, nil } -func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ time.Duration) error { +func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, retention time.Duration) error { m.mu.Lock() defer m.mu.Unlock() m.Commits++ @@ -173,6 +174,7 @@ func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ ti for _, c := range claims { if c.Status == dedupe.Claimed { m.committed[c.Key] = true + m.retention[c.Key] = retention delete(m.pending, c.Key) } } @@ -209,6 +211,13 @@ func (m *MockDeduplicator) Committed(k dedupe.Key) bool { return m.committed[k] } +// Retention is the retention k was last committed with. +func (m *MockDeduplicator) Retention(k dedupe.Key) time.Duration { + m.mu.Lock() + defer m.mu.Unlock() + return m.retention[k] +} + // Pending reports whether k is claimed and neither committed nor released. func (m *MockDeduplicator) Pending(k dedupe.Key) bool { m.mu.Lock() diff --git a/tests/e2e/fixtures/settings/config.json b/tests/e2e/fixtures/settings/config.json index a0d15cfc..a8d7c736 100644 --- a/tests/e2e/fixtures/settings/config.json +++ b/tests/e2e/fixtures/settings/config.json @@ -26,6 +26,7 @@ "enabled": true, "id_field": "event_id", "require_id": false, + "retention": "0", "tables": {} }, "dlq": { From a8e43bd2c41c285fa02d4b72e592e1b9f695ac82 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:13:48 -0400 Subject: [PATCH 030/108] fix(app): retry a failed DynamoDB table check; refuse no region A nested directory has no watcher, so a table check that failed at boot (a throttle, credentials not yet issued) left dedupe'd ingest failing closed until someone reloaded. The check now retries in the background, backing off 1s to 30s, and reconciles once it passes. NewDynamo refuses a config that resolves no region, a certain error caught at boot. Docs: reserve_concurrency has no effect while ingest sends one id per call; 0 = default; Pebble-only sentences scoped to dedupe.backend: pebble. Co-Authored-By: Claude Opus 5.5 (1M context) --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 6 +-- docs/src/content/docs/configuration.mdx | 14 +++---- docs/src/content/docs/deployment.md | 6 +-- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/dedupe_dynamodb_test.go | 41 ++++++++++++++++++++ internal/app/wire.go | 27 +++++++++++-- internal/config/backends.go | 4 +- internal/dedupe/dynamodb.go | 3 ++ 11 files changed, 87 insertions(+), 22 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 790a728c..4d062bbe 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (the in-process value by default; `dedupe.backend` also takes `dynamodb`, with its `dedupe.dynamodb` sub-block; `coord.backend` reserved) — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; conformance-tested against dynamodb-local, selected by `dedupe.backend: dynamodb`; boot checks the table and never creates it outside dynamodb-local), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; conformance-tested against dynamodb-local, selected by `dedupe.backend: dynamodb`; boot checks the table and never creates it outside dynamodb-local), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` or, gated on the table check (`Factory.Gated`), `Dynamo.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index f74c6cb5..8f6e31a4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/stores.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` Binary alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until a reload's check passes. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` Binary alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. diff --git a/config.yaml b/config.yaml index 7d4cbd45..93f66539 100644 --- a/config.yaml +++ b/config.yaml @@ -50,7 +50,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) lease: 30s # how long a claimed id stays pending; at most 2m with the embedded mq - reserve_concurrency: 64 # parallel calls per request to a remote backend + reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend; no effect yet (ingest sends one id per call) # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod # region: "" # empty = AWS_REGION diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 0bb4030c..a5ca5ca3 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with every reload checking again until it passes. It has no Pebble gauges. Both cases hand the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`). The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes: by a background component that backs off from one second to thirty (a nested directory has no watcher), and by every reload. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. One rule spans two layers: `dedupe.lease` may not exceed the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`) while `mq.backend` is `embedded`. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, selected by `dedupe.backend: dynamodb`: every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, selected by `dedupe.backend: dynamodb`: every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local, and refuses a config that resolves no region. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `Factory.Gated(ready)` wraps a factory so a store opens only once `ready` returns nil, and fails closed until then (the DynamoDB wiring's table check). `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 3a7b8b12..d6f071a0 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -35,7 +35,7 @@ This page is boot config only — what the platform operator owns (wiring, lifec | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble, holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | +| `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble (with `dedupe.backend: pebble`), holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | ### Backends @@ -56,20 +56,20 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`). | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most calls one request makes to a remote dedupe backend at once. `pebble` ignores it. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` = the default. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` = the default. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (Binary) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and every reload checks again. The check runs whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (Binary) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background (backing off from one second to thirty) and on every reload, so a table that comes good is picked up without a restart. The check runs whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(none)* | The shared table. Required. | -| `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`). | +| `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`); no region from either refuses boot. | | `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | -| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. | -| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. `0` = the default. | +| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. `0` = the default. | | `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. | | `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 5d0cb186..3919bec4 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -169,7 +169,7 @@ WH_SETTINGS_DIR=/etc/wavehouse/settings WaveHouse keeps all embedded state under a single configurable root, `WH_DATA_DIR` (yaml: `data_dir`). Subdirectories are convention, not config: - `/nats` — embedded NATS JetStream. Holds in-flight events between an ingest POST and the ingest worker → ClickHouse flush, plus the `stream.gap_window_minutes` window (settings directory) of history that powers SSE gap-fill across restarts. -- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). +- `/pebble` — the Pebble dedup KV (with `dedupe.backend: pebble`, the default): one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). In a Docker / Podman / Kubernetes deployment, **`data_dir` must resolve to a host-backed volume**. The reference compose file `deployments/compose/standalone.yaml` sets `WH_DATA_DIR=/app/data` and binds a `wavehouse-data:/app/data` volume — copy that pattern. The bundled Dockerfiles pre-create `/app/data` and `/app/settings` owned by the nonroot user (UID 65532); the binary creates the `nats/` and `pebble/` subdirectories under `/app/data` itself on first run. @@ -177,7 +177,7 @@ If `data_dir` resolves into the container's writable overlay layer instead, **Je Beyond persistence, the *speed* of that volume matters: JetStream `fsync`s every event to `/nats` before the ingest endpoint returns `200`, so the volume's `fsync` latency is your ingest latency floor. Managed cloud block storage handles this without thinking; commodity or virtualized substrates (ZFS without a SLOG, qcow2-on-`ext4`, spinning disks) can stall ingest with multi-second `fsync` tails. See [Durability & Storage](/durability) to measure yours before going live. -WaveHouse runs a simple existence check on startup and logs a `WARN` if `/nats` (or `/pebble`, when dedupe is on) is missing or empty: +WaveHouse runs a simple existence check on startup and logs a `WARN` if `/nats` (or `/pebble`, when dedupe is on with the `pebble` backend) is missing or empty: ```text wrap=false WARN data directory does not exist — starting with no prior state. @@ -493,7 +493,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and each reload checks the table again. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor `AWS_REGION`) refuses boot in both shapes. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 156421f7..c7377620 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: a table that fails it fails every tenant with dedupe on closed, whatever the switches say, until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index ded483a9..c86d2294 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -3,12 +3,14 @@ package app import ( "context" "io" + "net" "net/http" "net/http/httptest" "path/filepath" "strings" "sync" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -162,3 +164,42 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { require.NoError(t, err) }) } + +// A nested directory has no watcher, so a table that comes good is picked up +// by the background retry, not only by a reload someone has to send. +func TestRun_DynamoDBDedupeRetriesTheTableCheck(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + acme := a.dedup.For("acme") + require.False(t, acme.Open()) + + _, stop := runApp(t, a, ln) + fake.setExists(true) + require.Eventually(t, acme.Open, 10*time.Second, 50*time.Millisecond, "the retry opened the store without a reload") + require.NoError(t, stop()) +} + +// No region anywhere is a certain config error: refused at boot in either +// shape rather than failing every check afterwards. +func TestNew_DynamoDBDedupeRefusesNoRegion(t *testing.T) { + for name, dir := range map[string]func(*testing.T) string{ + "flat": func(t *testing.T) string { return writeSettings(t, dedupeOn) }, + "nested": func(t *testing.T) string { return writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn}) }, + } { + t.Run(name, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, dir(t)) + dynamoConfig(t, cfg, true) + cfg.Dedupe.DynamoDB.Region = "" + t.Setenv("AWS_REGION", "") + t.Setenv("AWS_DEFAULT_REGION", "") + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "dynamodb region is not set") + }) + } +} diff --git a/internal/app/wire.go b/internal/app/wire.go index 426db887..a2e23b7c 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -541,7 +541,10 @@ var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") // has dedupe on, and never creates it otherwise. A table that fails the check // follows the registry's rule for the shape, as Pebble's instance does: a // flat directory refuses boot; a nested one boots with every switched-on -// store closed, so its ingest fails closed, and each reload checks again. +// store closed, so its ingest fails closed. Unlike a local disk, a remote +// table's failure is usually brief (a throttle, credentials not yet issued +// mid-rollout), and a nested directory has no watcher to reload it, so the +// check is also retried in the background, with backoff, until it passes. func (a *App) wireDynamoDedupe(ctx context.Context) error { c := a.cfg.Dedupe.DynamoDB d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ @@ -576,7 +579,10 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) a.dedup = stores a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) + var reconciling sync.Mutex // the hook and the retry loop both reconcile reconcile := func(ctx context.Context) error { + reconciling.Lock() + defer reconciling.Unlock() if err := stores.Retain(a.served); err != nil { slog.Error("dedupe store close failed", "error", err) } @@ -598,8 +604,23 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { return checkErr } a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) - if err := reconcile(ctx); err != nil && !a.tenants.Nested() { - return fmt.Errorf("dedupe open: %w", err) + if err := reconcile(ctx); err != nil { + if !a.tenants.Nested() { + return fmt.Errorf("dedupe open: %w", err) + } + a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { + for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { + select { + case <-ctx.Done(): + return nil + case <-time.After(wait): + } + if reconcile(ctx) == nil { + slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + } + } + return nil + }}) } return nil } diff --git a/internal/config/backends.go b/internal/config/backends.go index 9b0cb57a..0b586d57 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -76,8 +76,8 @@ type Dedupe struct { // Lease is how long a claimed id stays pending while its record is // published; a claim its request never settles lapses after it. Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE" env-default:"30s"` - // ReserveConcurrency bounds the parallel calls one request makes to a - // remote backend. Pebble ignores it. + // ReserveConcurrency bounds the parallel calls one Reserve, Commit or + // Release makes to a remote backend. Pebble ignores it. ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY" env-default:"64"` DynamoDB DedupeDynamoDBConfig `yaml:"dynamodb"` } diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index eb1c8383..f07904af 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -139,6 +139,9 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load if err != nil { return nil, fmt.Errorf("dedupe: aws config: %w", err) } + if awsCfg.Region == "" { + return nil, errors.New("dedupe: dynamodb region is not set: set dedupe.dynamodb.region or AWS_REGION") + } client := dynamodb.NewFromConfig(awsCfg, func(o *dynamodb.Options) { if cfg.Endpoint != "" { o.BaseEndpoint = aws.String(cfg.Endpoint) From 6c74b65969b26e6cc5bc94ac6957f001ffbed8cb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:16:23 -0400 Subject: [PATCH 031/108] test(dedupe): pin that a commit mid-sweep-chunk survives; docs wording Adds a sweep hook so a Commit can race into the gap between a chunk's read and delete; the test fails (3/3) with commitMu removed. Docs: the sweep follows the shared instance, reads (not deletes) 1,024 keys per chunk, and "0" is the one unitless retention. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/durability.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/dedupe/embedded.go | 3 ++ internal/dedupe/sweep.go | 3 ++ internal/dedupe/sweep_test.go | 35 ++++++++++++++++++++ 6 files changed, 44 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f47f81ba..03affb5f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,7 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. -- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It deletes 1,024 keys per chunk without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index a0a62952..2ad53395 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads and deletes 1,024 keys at a time, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. +With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads 1,024 keys at a time, deleting the expired ones, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 126f35dd..b9d30162 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -190,7 +190,7 @@ Every dedupe knob lives here — there are no boot-config keys for it. The switc - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep deletes the expired id from the store: first a minute after the store opens, then hourly, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"` or a bare number. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. - `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 5186a738..1a620412 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -47,6 +47,9 @@ type Embedded struct { // readHook, when set, runs before each Pebble read in Reserve; a test // makes it fail to exercise Reserve's all-or-nothing error path. readHook func() error + // sweepHook, when set, runs in a sweep chunk between reading its keys and + // deleting them; a test races a Commit into that gap. + sweepHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index e536c5ed..396d6455 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -135,6 +135,9 @@ func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, r if err := it.Close(); err != nil { return nil, fmt.Errorf("dedupe sweep: %w", err) } + if e.sweepHook != nil { + e.sweepHook() + } // NoSync: a delete lost to a crash is redone by the next pass. if err := b.Commit(pebble.NoSync); err != nil { return nil, fmt.Errorf("dedupe sweep: %w", err) diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index 44579dcd..0589c103 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -97,6 +97,41 @@ func TestEmbedded_SweepDeletesExpiredAndVersionZeroKeys(t *testing.T) { assert.Equal(t, sweepResult{}, res, "a second pass finds nothing") } +// A Commit that arrives while a sweep chunk has read an expired key but not +// yet deleted it waits for the chunk, so the new commit is never deleted with +// the old value. Without the lock the Commit lands in the gap and the sweep +// then deletes it; the wait below only ever lets that pass, never fail. +func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + clock := newStepClock() + SetClock(e, clock.now) + m := switchedOn(t, e, "acme") + commitIDs(t, m, time.Hour, "e1") + clock.advance(2 * time.Hour) + + claims, err := m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) + require.NoError(t, err) + require.Equal(t, Claimed, claims[0].Status, "expired: claimable again") + done := make(chan error, 1) + e.sweepHook = func() { + go func() { done <- m.Commit(context.Background(), claims, time.Hour) }() + select { + case <-done: + done <- nil + case <-time.After(50 * time.Millisecond): + } + } + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Expired: 1}, res) + require.NoError(t, <-done) + + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.True(t, dup, "the commit made mid-chunk survived the sweep") +} + // A retention is honoured on read before any sweep has run: the key is a // duplicate until the retention ends and claimable from that instant. func TestEmbedded_RetentionHonouredOnRead(t *testing.T) { From 0b604510236fa0eb19dea64a2fcde189f220452b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:20:25 -0400 Subject: [PATCH 032/108] docs(dedupe): name every region source; untangle the dynamodb check clause Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 3919bec4..b889008a 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -493,7 +493,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor `AWS_REGION`) refuses boot in both shapes. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c7377620..524c39dc 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: a table that fails it fails every tenant with dedupe on closed, whatever the switches say, until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. From fc4085b1a6f4c20b85a0369b3e4bd6cc51617b37 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 11:25:00 -0400 Subject: [PATCH 033/108] docs(durability): describe the benchmark machine neutrally Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/durability.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 5a736b67..dc6ff7f3 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -60,7 +60,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) ## Deduplication: one more fsync per window -With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. From 02d7056bf76c42aa712758aaffa718a4785bfed5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 11:25:11 -0400 Subject: [PATCH 034/108] docs(deployment): use neutral tags in the DynamoDB Terraform example Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/deployment.md | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 9bea802e..c7e88c2c 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -440,7 +440,7 @@ What the backend requires of the table: Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. -An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: +An example in Terraform. Replace the tags with your own conventions: ```hcl resource "aws_dynamodb_table" "wavehouse_dedupe" { @@ -465,10 +465,9 @@ resource "aws_dynamodb_table" "wavehouse_dedupe" { tags = { Name = "wavehouse-dedupe-${var.environment}" - Project = "wavehouse-cloud" - Environment = var.environment # prod | dev | ci | demo | benchmark - ManagedBy = "wavehouse-cloud/infra/stacks/prod-platform" - CostCenter = "data-plane" + Project = "wavehouse" + Environment = var.environment + ManagedBy = "terraform" } } From d3032cc1a10984d9cebe6dc7ae553430e63f98f9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:22:33 -0400 Subject: [PATCH 035/108] fix(dedupe): length-prefix the key's fields; refuse no table name MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The key separated tenant, table and id with NUL, so a table name holding NUL had to be refused. Each field before the id now carries its length as a uvarint instead: 0x01 ‖ uvarint(len(tenant)) ‖ tenant ‖ uvarint(len(table)) ‖ table ‖ id Every key parses back to the one triple that wrote it, so any table name ClickHouse accepts dedupes in a keyspace of its own. The id stays last and unprefixed, so long-id hashing is unchanged. The version byte stays 0x01: the NUL layout was never released, and the earlier layouts (a bare id in v0.1.0, tenant ‖ 0x00 ‖ id after it) store 8-byte values a version-1 read treats as absent. Settings validation and Key.Validate no longer refuse NUL; with nothing left to refuse, Key.Validate and ErrInvalidKey are removed. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dedupetest/dedupetest.go | 31 +++++-- internal/dedupe/key.go | 46 ++++------ internal/dedupe/key_layout_test.go | 103 +++++++++++++++++++++++ internal/dedupe/managed.go | 12 +-- internal/settings/validate.go | 8 +- internal/settings/validate_test.go | 15 +++- 8 files changed, 168 insertions(+), 51 deletions(-) create mode 100644 internal/dedupe/key_layout_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index 56aa5a3a..646f3614 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,7 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so no table name is refused: one holding a NUL byte, or any other, dedupes in a keyspace of its own. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 74066932..dba5ffb7 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -123,7 +123,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `dedupe/` — Deduplication (Optional) - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. -- **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. +- **key.go** — the key layout every backend stores: a version byte, the tenant's length and the tenant, the table's length and the table, then the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The lengths (uvarints) rather than a separator mark where each field ends, so a table name may hold any byte — NUL included — and neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index c3d1cd9c..a9934001 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -6,7 +6,6 @@ package dedupetest import ( "context" - "errors" "fmt" "strings" "sync" @@ -308,11 +307,31 @@ var cases = []struct { require.NoError(t, d.Commit(t.Context(), dup, 0), "only Claimed claims commit") assert.Equal(t, dedupe.Claimed, reserve(t, d, long, key("e1"))[0].Status) }}, - {"a table holding NUL is refused", func(t *testing.T, s *suite) { + {"any table name is its own keyspace", func(t *testing.T, s *suite) { d := s.store(t, "acme") - _, err := d.Reserve(t.Context(), []dedupe.Key{key("ok"), {Table: "a\x00b", ID: "e1"}}, long) - require.ErrorIs(t, err, dedupe.ErrInvalidKey) - assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a bad key is not worth retrying") - assert.Equal(t, dedupe.Claimed, reserve(t, d, long, key("ok"))[0].Status, "and nothing was claimed") + // seen[i] and fresh[i] differ only in where table ends and id + // begins; the first two pairs would share one key under a + // NUL-separated layout. + seen := []dedupe.Key{ + {Table: "a", ID: "b\x00c"}, + {Table: "a\x00", ID: "b"}, + {Table: "", ID: "\x01a"}, + {Table: "\xff\xfe", ID: "e1"}, + {Table: "tab\tle \n", ID: "e1"}, + } + fresh := []dedupe.Key{ + {Table: "a\x00b", ID: "c"}, + {Table: "a", ID: "\x00b"}, + {Table: "\x01", ID: "a"}, + {Table: "\xff", ID: "\xfee1"}, + {Table: "tab\tle", ID: " \ne1"}, + } + require.NoError(t, d.Commit(t.Context(), reserve(t, d, long, seen...), 0)) + for _, c := range reserve(t, d, long, fresh...) { + assert.Equal(t, dedupe.Claimed, c.Status, "%q", c.Key) + } + for _, c := range reserve(t, d, long, seen...) { + assert.Equal(t, dedupe.Duplicate, c.Status, "%q", c.Key) + } }}, } diff --git a/internal/dedupe/key.go b/internal/dedupe/key.go index 25ead7c2..c1cfc97e 100644 --- a/internal/dedupe/key.go +++ b/internal/dedupe/key.go @@ -2,24 +2,23 @@ package dedupe import ( "crypto/sha256" - "errors" - "fmt" - "strings" + "encoding/binary" "github.com/Wave-RF/WaveHouse/internal/tenant" ) // The key layout every backend stores, byte for byte: // -// keyVersion ‖ tenant ‖ keySeparator ‖ table ‖ keySeparator ‖ id +// keyVersion ‖ uvarint(len(tenant)) ‖ tenant ‖ uvarint(len(table)) ‖ table ‖ id // -// A tenant id is letters, digits, '_' and '-', so it never holds the -// separator and never starts with keyVersion — the version-0 keys before -// #222 (tenant ‖ 0x00 ‖ id) never meet these. The id is last, so it may hold -// anything. +// Each field before the id carries its length, so a table name may hold any +// byte — NUL included — and no two (tenant, table, id) triples share a key. +// The id is last, so it needs no length and may hold anything too. A tenant +// id never starts with keyVersion (tenant.Parse admits letters, digits, '_' +// and '-'), so the version-0 keys before #222 (tenant ‖ 0x00 ‖ id) never meet +// these. const ( - keyVersion byte = 0x01 - keySeparator byte = 0x00 + keyVersion byte = 0x01 // hashedID leads an id stored as its SHA-256 rather than verbatim. Ids // that start with it are hashed too, so a verbatim id never reads as a // hashed one. @@ -29,25 +28,12 @@ const ( MaxIDBytes = 1024 ) -// ErrInvalidKey is returned for a key no backend can store: a table name -// holding the separator byte. -var ErrInvalidKey = errors.New("invalid dedupe key") - // KeyPrefix is the part of every key that names tenant id, so a backend // computes it once per tenant store. func KeyPrefix(id tenant.ID) []byte { - p := make([]byte, 0, len(id)+2) + p := make([]byte, 0, len(id)+1+binary.MaxVarintLen64) p = append(p, keyVersion) - p = append(p, id...) - return append(p, keySeparator) -} - -// Validate reports whether k can be stored. -func (k Key) Validate() error { - if strings.IndexByte(k.Table, keySeparator) >= 0 { - return fmt.Errorf("%w: table name %q holds a NUL byte", ErrInvalidKey, k.Table) - } - return nil + return appendField(p, string(id)) } // Hashed reports whether k's id is stored as its SHA-256 rather than @@ -57,11 +43,10 @@ func (k Key) Hashed() bool { } // AppendKey appends k's stored form, under the tenant prefix from KeyPrefix, -// to dst. k must be valid. +// to dst. func AppendKey(dst, prefix []byte, k Key) []byte { dst = append(dst, prefix...) - dst = append(dst, k.Table...) - dst = append(dst, keySeparator) + dst = appendField(dst, k.Table) if k.Hashed() { sum := sha256.Sum256([]byte(k.ID)) dst = append(dst, hashedID) @@ -69,3 +54,8 @@ func AppendKey(dst, prefix []byte, k Key) []byte { } return append(dst, k.ID...) } + +func appendField(dst []byte, s string) []byte { + dst = binary.AppendUvarint(dst, uint64(len(s))) + return append(dst, s...) +} diff --git a/internal/dedupe/key_layout_test.go b/internal/dedupe/key_layout_test.go new file mode 100644 index 00000000..224382ba --- /dev/null +++ b/internal/dedupe/key_layout_test.go @@ -0,0 +1,103 @@ +package dedupe_test + +import ( + "crypto/sha256" + "encoding/binary" + "strings" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// The layout is pinned byte for byte: DynamoDB items and Pebble keys outlive +// the binary that wrote them. +func TestKeyLayout(t *testing.T) { + t.Parallel() + got := dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "a\x00b", ID: "e1"}) + assert.Equal(t, []byte("\x01\x04acme\x03a\x00be1"), got) + + long := strings.Repeat("x", dedupe.MaxIDBytes+1) + sum := sha256.Sum256([]byte(long)) + got = dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "t", ID: long}) + assert.Equal(t, append([]byte("\x01\x04acme\x01t\xff"), sum[:]...), got) + + table := strings.Repeat("t", 200) + got = dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: table, ID: "e1"}) + assert.Equal(t, []byte("\x01\x04acme\xc8\x01"+table+"e1"), got, "a length past 127 takes two bytes") +} + +// decodeKey inverts AppendKey. That it exists — every key parses back to the +// one triple that wrote it — is what makes the layout collision-free. +func decodeKey(t *testing.T, b []byte) (tn, table, idPart string) { + t.Helper() + require.NotEmpty(t, b) + require.Equal(t, byte(0x01), b[0]) + b = b[1:] + field := func() string { + n, w := binary.Uvarint(b) + require.Positive(t, w) + b = b[w:] + require.LessOrEqual(t, n, uint64(len(b))) + s := string(b[:n]) + b = b[n:] + return s + } + tn = field() + table = field() + return tn, table, string(b) +} + +// Triples a separator could confuse — NUL or the version byte in the table, +// in the id, at either end, or moved across the table/id boundary — each get +// a key of their own, and parse back to themselves. +func TestKeyLayout_NoCollisions(t *testing.T) { + t.Parallel() + tenants := []tenant.ID{"a", "ab", "a_b", "acme"} + pieces := []string{"", "\x00", "\x01", "\xff", "a", "b", "a\x00", "\x00b", "a\x00b", "\x01\x04acme", "\x03a"} + seen := map[string]string{} + check := func(tn tenant.ID, k dedupe.Key) { + key := dedupe.AppendKey(nil, dedupe.KeyPrefix(tn), k) + who := string(tn) + " / " + k.Table + " / " + k.ID + if prev, ok := seen[string(key)]; ok && prev != who { + t.Fatalf("%q and %q share the key %q", prev, who, key) + } + seen[string(key)] = who + gotTenant, gotTable, idPart := decodeKey(t, key) + assert.Equal(t, string(tn), gotTenant) + assert.Equal(t, k.Table, gotTable) + if !k.Hashed() { + assert.Equal(t, k.ID, idPart) + } + } + for _, tn := range tenants { + for _, table := range pieces { + for _, id := range pieces { + check(tn, dedupe.Key{Table: table, ID: id}) + } + } + } + // Every table and id up to three bytes over an alphabet of the bytes a + // layout could misread. + var all []string + var grow func(prefix string) + grow = func(prefix string) { + all = append(all, prefix) + if len(prefix) < 3 { + for _, c := range []string{"\x00", "\x01", "\x04", "\xff", "a"} { + grow(prefix + c) + } + } + } + grow("") + for _, tn := range tenants { + for _, table := range all { + for _, id := range all { + check(tn, dedupe.Key{Table: table, ID: id}) + } + } + } +} diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index 041487cb..0c53f97d 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -84,16 +84,12 @@ func (m *Managed) Open() bool { return m.db != nil } -// Reserve checks every key is storable, collapses a key repeated inside keys -// to one backend claim — later occurrences answer Duplicate — reads a lease -// <= 0 as DefaultLease, and delegates the rest to the open store; -// ErrDisabled while switched off, ErrUnavailable while switched on but not -// open. +// Reserve collapses a key repeated inside keys to one backend claim — later +// occurrences answer Duplicate — reads a lease <= 0 as DefaultLease, and +// delegates the rest to the open store; ErrDisabled while switched off, +// ErrUnavailable while switched on but not open. func (m *Managed) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { for _, k := range keys { - if err := k.Validate(); err != nil { - return nil, err - } if k.Hashed() { hashedIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", k.Table))) } diff --git a/internal/settings/validate.go b/internal/settings/validate.go index 08575c95..598c4aa7 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -451,18 +451,14 @@ func (v *validator) checkIDField(path string, val *string) { } // checkTableName rejects a per-table override key that could never match a -// table: empty, carrying surrounding whitespace, or holding NUL. Shared by -// the dedupe and dlq override maps. +// table: empty, or carrying surrounding whitespace. Shared by the dedupe and +// dlq override maps. func (v *validator) checkTableName(mapPath, table string) { switch { case table == "": v.errorf(FileConfig, mapPath, "table name must not be empty") case strings.TrimSpace(table) != table: v.errorf(FileConfig, mapPath+"."+table, "table name %q has surrounding whitespace", table) - case strings.ContainsRune(table, 0): - // A NUL separates the dedupe key's fields (dedupe.AppendKey), so a - // table holding one could never be deduped. - v.errorf(FileConfig, mapPath+"."+table, "table name %q holds a NUL byte", table) } } diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index 4dae9c9d..d9dd3479 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -264,7 +264,6 @@ func TestValidate_ContentRules(t *testing.T) { {"padded override id_field", FileConfig, `{"dedupe": {"tables": {"clicks": {"id_field": "click_id "}}}}`, "dedupe.tables.clicks.id_field"}, {"empty override table name", FileConfig, `{"dedupe": {"tables": {"": {"id_field": "x"}}}}`, "table name must not be empty"}, {"override table whitespace", FileConfig, `{"dedupe": {"tables": {" clicks": {"require_id": true}}}}`, "surrounding whitespace"}, - {"override table NUL", FileConfig, `{"dedupe": {"tables": {"cli\u0000cks": {"require_id": true}}}}`, "holds a NUL byte"}, {"empty override id_field", FileConfig, `{"dedupe": {"tables": {"clicks": {"id_field": ""}}}}`, "dedupe.tables.clicks.id_field: must not be empty"}, {"negative max rows", FileConfig, `{"query": {"default_max_rows": -1}}`, "must be >= 1"}, {"zero max rows", FileConfig, `{"query": {"default_max_rows": 0}}`, "must be >= 1"}, @@ -344,6 +343,20 @@ func TestValidate_ContentRules(t *testing.T) { } } +// A table name with odd bytes — NUL included — is any other table name to +// the override maps: dedupe keys carry the table's length, not a separator. +func TestValidate_OverrideTableNamesAnyBytes(t *testing.T) { + t.Parallel() + files := validFiles() + files[FileConfig] = configJSON(`{"dedupe": {"tables": {"cli\u0000cks": {"require_id": true}, "a\u0001b\tc": {"id_field": "x"}}}, "dlq": {"tables": {"cli\u0000cks": {"enabled": false}}}}`) + doc, findings := ValidateDir(writeDir(t, files)) + require.Empty(t, findings, "findings: %s", findingStrings(findings)) + require.NotNil(t, doc) + assert.Contains(t, doc.Config.Dedupe.Tables, "cli\x00cks") + assert.Contains(t, doc.Config.Dedupe.Tables, "a\x01b\tc") + assert.Contains(t, doc.Config.DLQ.Tables, "cli\x00cks") +} + // TestValidate_ClickHouseTLSPathsAreNotOpened pins that the tls block is // checked for shape only: Validate is pure and also runs on the control // plane, so paths that exist nowhere still validate, and the values reach From 7b9595e35cd9a5d2aca3ab5ef50cd7527839f5ec Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:25:03 -0400 Subject: [PATCH 036/108] feat(settings): a missing dedupe.retention keeps ids forever MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit dedupe.retention was required in every config.json. It is now optional: missing means "0" (forever) at the tenant level, and a table override without it inherits the tenant's, as for the other override fields. A present value is validated as before — unparseable, negative, or finite below the queue's duplicate window is refused — so an existing directory needs no change to upgrade. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 8 +++---- internal/settings/settings.go | 11 +++++---- internal/settings/store.go | 5 +++- internal/settings/store_test.go | 17 +++++++++++++ internal/settings/validate.go | 7 ++---- internal/settings/validate_test.go | 25 +++++++++++++++++++- 8 files changed, 59 insertions(+), 18 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 03affb5f..ffa7cb6f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,7 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. -- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains an optional **`dedupe.retention`** key, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. A `config.json` without the key keeps ids forever, and a table override without one inherits the tenant's, so an existing directory needs no change. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 86347f44..ca23f0fe 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -421,7 +421,7 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys are never read, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. -The same release adds **`dedupe.retention`, a required key**: every `config.json`, each tenant's folder included, must state it or the directory is refused (at boot) or not adopted (on reload). `"retention": "0"` keeps every id forever, as before; see [Deduplication](/settings-directory#deduplication) for a finite one. +The same release adds an optional **`dedupe.retention`** key. No upgrade step is needed: a `config.json` without it keeps every id forever, as before. See [Deduplication](/settings-directory#deduplication) for a finite one. ## Upgrading across the v2 ingest envelope diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index b9d30162..55a8e0d9 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -99,7 +99,7 @@ Pipes are read per request, so a reload changes what the next `GET /v1/pipes/{na ## `config.json` keys -The tenant tunables. Every key is required (a missing one is a validation error) except the per-table overrides; the "Seed" column is what `wavehouse bootstrap` writes: +The tenant tunables. Every key is required (a missing one is a validation error) except `dedupe.retention` and the per-table overrides; the "Seed" column is what `wavehouse bootstrap` writes: | Key | Seed | Description | | --- | ---- | ----------- | @@ -123,7 +123,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `dedupe.enabled` | `false` | Turn deduplication on; a reload opens or closes this tenant's store — see [Deduplication](#deduplication). | | `dedupe.id_field` | `event_id` | Dedup key field — see [Deduplication](#deduplication). | | `dedupe.require_id` | `false` | Reject rows missing the id field — see [Deduplication](#deduplication). | -| `dedupe.retention` | `"0"` | How long a committed id stays a duplicate, as a duration (`"720h"`); `"0"` keeps it forever — see [Deduplication](#deduplication). | +| `dedupe.retention` | `"0"` | How long a committed id stays a duplicate, as a duration (`"720h"`); `"0"`, or leaving the key out, keeps it forever — see [Deduplication](#deduplication). | | `dedupe.tables.
.{id_field, require_id, retention}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | | `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the tenant's dead-letter stream (`DLQ_{tenant}`) (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | | `dlq.tables.
.enabled` | `{}` | Optional per-table override of the switch. | @@ -190,8 +190,8 @@ Every dedupe knob lives here — there are no boot-config keys for it. The switc - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. -- `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. +- `dedupe.retention` (optional; seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed, and a `config.json` without the key means the same. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest, so a table with no `retention` keeps the tenant's (forever when the tenant sets none). A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 68445543..eeb03f35 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -143,8 +143,9 @@ type AuthConfig struct { // (dedupe.Managed, one per tenant, each a share of the one embedded Pebble // instance), so the whole block is tenant-owned. // -// id_field, require_id and retention are required here and optional per -// table: a table override inherits whichever field it doesn't name. An empty, +// id_field and require_id are required here; retention is optional, and +// missing means "0" (forever). Every field is optional per table: a table +// override inherits whichever field it doesn't name. An empty, // whitespace-only, or whitespace-padded id_field is rejected at every level, // so the effective id_field can never be empty or silently unmatchable. type DedupeConfig struct { @@ -152,9 +153,9 @@ type DedupeConfig struct { IDField *string `json:"id_field"` RequireID *bool `json:"require_id"` // Retention is how long a committed id stays a duplicate, as a Go - // duration ("720h"); "0" keeps it forever. A change applies to ids - // committed after it. - Retention *string `json:"retention"` + // duration ("720h"); "0", or leaving it out, keeps it forever. A change + // applies to ids committed after it. + Retention *string `json:"retention,omitempty"` // Tables holds per-table overrides keyed by ClickHouse table name (#222). // Names are format-checked only — existence is schema discovery's runtime // concern, same as policies.json table keys. diff --git a/internal/settings/store.go b/internal/settings/store.go index b3615efe..d707451d 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -93,7 +93,10 @@ type Dedupe struct { func (s *Store) DedupeFor(table string) Dedupe { d := s.doc().Config.Dedupe out := Dedupe{Enabled: *d.Enabled, IDField: *d.IDField, RequireID: *d.RequireID} - retention := *d.Retention + retention := "0" + if d.Retention != nil { + retention = *d.Retention + } if td, ok := d.Tables[table]; ok { if td.IDField != nil { out.IDField = *td.IDField diff --git a/internal/settings/store_test.go b/internal/settings/store_test.go index e9689824..3487ba50 100644 --- a/internal/settings/store_test.go +++ b/internal/settings/store_test.go @@ -1,6 +1,7 @@ package settings import ( + "encoding/json" "os" "path/filepath" "testing" @@ -66,6 +67,22 @@ func TestStore_DedupeFor_Cascade(t *testing.T) { } } +// A config.json without dedupe.retention keeps ids forever, and its table +// overrides inherit that or set their own. +func TestStore_DedupeFor_RetentionMissing(t *testing.T) { + t.Parallel() + var doc map[string]map[string]any + require.NoError(t, json.Unmarshal([]byte(configJSON(`{"dedupe": {"tables": {"clicks": {"id_field": "click_id"}, "views": {"retention": "24h"}}}}`)), &doc)) + delete(doc["dedupe"], "retention") + body, err := json.Marshal(doc) + require.NoError(t, err) + s := newLoadedStore(t, map[string]string{FileConfig: string(body)}) + + assert.Equal(t, Dedupe{IDField: "event_id"}, s.DedupeFor("other"), "forever") + assert.Equal(t, Dedupe{IDField: "click_id"}, s.DedupeFor("clicks"), "inherits forever") + assert.Equal(t, Dedupe{IDField: "event_id", Retention: 24 * time.Hour}, s.DedupeFor("views")) +} + // TestStore_SeedIsValid pins that the shipped starter directory passes its // own gate: `wavehouse bootstrap` must never write something // `wavehouse validate` rejects, and the defaults are readable back. diff --git a/internal/settings/validate.go b/internal/settings/validate.go index 76352e19..52fd7637 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -454,8 +454,8 @@ func (v *validator) checkIDField(path string, val *string) { // checkRetention rejects a dedupe retention that is not a duration, is // negative, or is finite but shorter than MinDedupeRetention. The short one // is refused rather than raised to the minimum, so the file never means -// something other than what it says. nil is the caller's concern, as for -// id_field. +// something other than what it says. nil is valid: forever at the tenant +// level, inherited at the table level. func (v *validator) checkRetention(path string, val *string) { if val == nil { return @@ -694,9 +694,6 @@ func (v *validator) parseConfig(data []byte) TenantConfig { if d.RequireID == nil { v.required("dedupe.require_id") } - if d.Retention == nil { - v.required("dedupe.retention") - } v.checkIDField("dedupe.id_field", d.IDField) v.checkRetention("dedupe.retention", d.Retention) // Sorted iteration keeps finding order deterministic across runs. diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index a6d7c5f7..04894acc 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -285,7 +285,6 @@ func TestValidate_ContentRules(t *testing.T) { {"keepalive as a duration string", FileConfig, `{"stream": {"keepalive_interval": "30s"}}`, "keepalive_interval"}, {"missing dedupe.require_id", FileConfig, `{"dedupe": {"id_field": "event_id"}}`, "dedupe.require_id: required"}, {"missing dedupe.enabled", FileConfig, `{"dedupe": {"id_field": "event_id", "require_id": false, "retention": "0"}}`, "dedupe.enabled: required"}, - {"missing dedupe.retention", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}}`, "dedupe.retention: required"}, {"dedupe.retention not a duration", FileConfig, configJSON(`{"dedupe": {"retention": "30d"}}`), `dedupe.retention: must be a duration such as "720h"`}, {"dedupe.retention a number", FileConfig, configJSON(`{"dedupe": {"retention": 3600}}`), "retention"}, {"dedupe.retention negative", FileConfig, configJSON(`{"dedupe": {"retention": "-1h"}}`), "dedupe.retention: must not be negative"}, @@ -373,6 +372,30 @@ func TestValidate_DedupeRetentionAccepted(t *testing.T) { } } +// configJSONWithout is the seed config.json less dedupe.retention. +func configJSONWithout(t *testing.T) string { + t.Helper() + var doc map[string]map[string]json.RawMessage + require.NoError(t, json.Unmarshal([]byte(configJSON(`{}`)), &doc)) + delete(doc["dedupe"], "retention") + out, err := json.Marshal(doc) + require.NoError(t, err) + return string(out) +} + +// dedupe.retention may be left out: the tenant keeps ids forever, and a +// table override may still set one. +func TestValidate_DedupeRetentionOptional(t *testing.T) { + t.Parallel() + files := validFiles() + files[FileConfig] = configJSONWithout(t) + require.NotContains(t, files[FileConfig], "retention") + doc, findings := ValidateDir(writeDir(t, files)) + require.NotNil(t, doc, "findings: %s", findingStrings(findings)) + assert.False(t, HasErrors(findings), "findings: %s", findingStrings(findings)) + assert.Nil(t, doc.Config.Dedupe.Retention) +} + // TestValidate_ClickHouseTLSPathsAreNotOpened pins that the tls block is // checked for shape only: Validate is pure and also runs on the control // plane, so paths that exist nowhere still validate, and the values reach From 12741f8c1a7dcd009ebc427e71645bbb565459e6 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:25:37 -0400 Subject: [PATCH 037/108] docs(dedupe): Reserve refuses no key; scope the table-name CHANGELOG line Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 646f3614..40c34d81 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,7 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so no table name is refused: one holding a NUL byte, or any other, dedupes in a keyspace of its own. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so a table name is no longer refused for the bytes it holds: one holding a NUL byte dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index dba5ffb7..703b6ec0 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -125,7 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant's length and the tenant, the table's length and the table, then the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The lengths (uvarints) rather than a separator mark where each field ends, so a table name may hold any byte — NUL included — and neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. +- **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. From 76b6a9e23d8055db8e6dc9c76b34b4e0e8cc8c15 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:32:49 -0400 Subject: [PATCH 038/108] docs(settings): name dedupe.retention as the one compiled default Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- config.yaml | 4 ++-- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/api/settings_test.go | 2 +- internal/app/wire.go | 3 ++- internal/settings/settings.go | 5 +++-- internal/settings/store.go | 6 +++--- 8 files changed, 14 insertions(+), 12 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b81b9031..8806ebbb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, deleted by the retention sweep below ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so a table name is no longer refused for the bytes it holds: one holding a NUL byte dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, deleted by the retention sweep (see Added) ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so a table name is no longer refused for the bytes it holds: one holding a NUL byte dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/config.yaml b/config.yaml index 35812417..8675e2de 100644 --- a/config.yaml +++ b/config.yaml @@ -67,8 +67,8 @@ auth: # query.default_max_rows / timestamp_bucket_seconds, # schema.refresh_interval, stream keepalive_interval / keepalive_buckets / # gap_window_minutes, mq.max_bytes_gb, cors.allowed_origins — and every key -# is required: the -# binary has no compiled defaults, so what's adopted is exactly what the +# is required except dedupe.retention (missing = "0", forever): the binary +# has no other compiled default, so what's adopted is exactly what the # files say. The server validates the directory at boot (invalid or missing # refuses to start) and reloads it on SIGHUP, on file change, or via # POST /v1/ops/settings/reload; a reload that fails validation keeps the diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e9a581e8..6ab56b3d 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -188,7 +188,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi - **store.go** — `Store` is a passive holder: one tenant's adopted document behind an atomic pointer, swapped by the registry, which stamps it with the id of the tenant it created the store for (`Tenant()`, how a handler names its tenant to a per-tenant resource). Consumers read typed accessors per call (`ClickHouse()`, `Auth()`, `DedupeFor(table)`, `DLQFor(table)`, `Keepalive()`, …) rather than holding values. - **registry.go** — `Registry` maps a tenant id to its `Store` and owns everything that changes one. `Open` validates and adopts at boot; `Reload` re-validates the whole directory and `ReloadTenant` one tenant's folder, serialized with each other; `AfterAdopt` hooks run after every reload the registry applied, with the tenants it adopted — none when it only rejected or removed one, which a consumer holding a resource per tenant needs to hear of too; `For(id)` and `All()` see only the tenants being served, `Known()` every tenant held, rejected ones included, and `Resolve(id)` tells a rejected tenant from an unknown one. The shape is fixed at `Open`. Flat: an invalid directory refuses boot, and a rejected reload keeps the previous snapshot. Nested: fail closed per tenant — a folder with an error finding stops being served (the store keeps its document for requests already admitted, and gets the next good one) while the rest carry on; a whole-directory reload mirrors the folders, down to none (an emptied directory is not a change of shape); and a finding about the directory itself refuses boot or rejects the reload whole, leaving every tenant as it was. The tenant map is replaced whole by a reload, so a lookup is one lock-free load. - **watch.go** — `Registry.Watch`, which `internal/app` starts for a flat directory only: fsnotify on the *directory* (not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost), debounced into one reload; reloads once as soon as the watch exists so an edit between the boot read and the watch is never missed. `SIGHUP` and the reload endpoint funnel through the same serialized `Reload`. -- **seed.go** / **seed/** — The embedded (`go:embed`) starter directory with every key at its default. The binary carries no compiled defaults: `wavehouse bootstrap [dir]` writes this seed, and the compose stack and e2e fixture ship copies of it. +- **seed.go** / **seed/** — The embedded (`go:embed`) starter directory with every key at its default. The binary carries no compiled defaults except that a missing `dedupe.retention` means `"0"`: `wavehouse bootstrap [dir]` writes this seed, and the compose stack and e2e fixture ship copies of it. ### `tenant/` — Tenant Identifier diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 55a8e0d9..9cb52b4d 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -11,7 +11,7 @@ Boot config — the YAML file and `WH_*` environment variables on the [Configura The settings directory holds WaveHouse's file-based settings as exactly four JSON documents: [`roles.json`](#rolesjson), [`policies.json`](#policiesjson), [`pipes.json`](#pipesjson), and [`config.json`](#configjson-keys). The files are the only write path — standalone, you edit them on the host; on WaveHouse Cloud the control plane writes them — and there is no API that writes back to them. Every file must exist (an empty document is `{}` — a missing file always means deletion or a wrong path, never "defaults"), and any other entry in the directory is an error, so a typoed filename or a stray backup fails loudly instead of being silently ignored. Dot-prefixed entries are the one exception: editor swap files and the `..data` machinery Kubernetes ConfigMap mounts publish through are ignored. -Create one with `wavehouse bootstrap [dir]`: it writes all four files with every key at its default and refuses a non-empty directory, so an existing settings directory is never overwritten. The binary carries no compiled defaults — the seed is the one place they live, and what the server adopts is exactly what the files say. The seed ships no policy (`policies.json` is `{}`, `roles.json` and `pipes.json` are empty lists), so a freshly bootstrapped directory boots fail-closed — every request is denied until you write a policy. The container images ship no settings directory: they preset `WH_SETTINGS_DIR=/app/settings` and expect a bind mount there — a host directory you wrote with `bootstrap` (the reference compose file mounts the checked-in `deployments/compose/settings/`, the seed with `clickhouse.addr` pointed at the `clickhouse` service and a permissive `public` trial policy in `policies.json` / `roles.json`). A bind mount, not a named volume: the images are distroless, with no shell to edit files inside a volume. A missing mount refuses to boot rather than running on defaults nobody chose. +Create one with `wavehouse bootstrap [dir]`: it writes all four files with every key at its default and refuses a non-empty directory, so an existing settings directory is never overwritten. The binary carries no compiled defaults but one — a missing `dedupe.retention` means `"0"`, forever — so the seed is where the defaults live, and what the server adopts is exactly what the files say. The seed ships no policy (`policies.json` is `{}`, `roles.json` and `pipes.json` are empty lists), so a freshly bootstrapped directory boots fail-closed — every request is denied until you write a policy. The container images ship no settings directory: they preset `WH_SETTINGS_DIR=/app/settings` and expect a bind mount there — a host directory you wrote with `bootstrap` (the reference compose file mounts the checked-in `deployments/compose/settings/`, the seed with `clickhouse.addr` pointed at the `clickhouse` service and a permissive `public` trial policy in `policies.json` / `roles.json`). A bind mount, not a named volume: the images are distroless, with no shell to edit files inside a volume. A missing mount refuses to boot rather than running on defaults nobody chose. Check a directory with `wavehouse validate [dir]`. Both commands resolve the directory the same way — the argument, falling back to `WH_SETTINGS_DIR`, and a usage error (exit `2`) with neither — so the path you seed is the path you validate, and inside the container images (which preset `WH_SETTINGS_DIR=/app/settings`) both work with no argument at all. `validate` validates without starting the server (JSON syntax including unknown fields and duplicate keys, per-file shape rules including the required keys, and cross-file role references), prints every finding in one pass, and exits `0` for valid (warnings allowed), `1` for invalid, `2` for usage — so operators and CI can gate a settings change before it reaches a running instance. diff --git a/internal/api/settings_test.go b/internal/api/settings_test.go index 24fa80f1..462cbc8e 100644 --- a/internal/api/settings_test.go +++ b/internal/api/settings_test.go @@ -16,7 +16,7 @@ import ( "github.com/stretchr/testify/require" ) -// fullConfig is a complete config.json (every key is required) with the +// fullConfig is a complete config.json (every key set) with the // given query.default_max_rows. func fullConfig(maxRows int) string { return fmt.Sprintf(`{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": %d, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, maxRows) diff --git a/internal/app/wire.go b/internal/app/wire.go index ac494bea..2678646a 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -52,7 +52,8 @@ func withoutContext(release func() error) func(context.Context) error { // configuration (dedupe, dlq, query, schema, stream, cors — see // settings.TenantConfig). Required: config.Validate already rejected an // empty settings.dir, and an invalid directory refuses boot. The binary -// carries no compiled defaults; `wavehouse bootstrap` writes the seed. A +// carries no compiled defaults but a missing dedupe.retention ("0"); +// `wavehouse bootstrap` writes the seed. A // *reload* of an invalid directory merely keeps the previous snapshot. A // nested directory (one folder per tenant, #583) fails closed per tenant // instead, at boot and on reload alike: see settings.Registry. diff --git a/internal/settings/settings.go b/internal/settings/settings.go index eeb03f35..4d5fa7a6 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -65,8 +65,9 @@ type PipesFile struct { // `cache.l1_max_cost`), listeners, the observability // exporters — and the secrets (`clickhouse.password`, `auth.jwt_secret`, // `auth.operator_key`), which never belong in a tracked JSON file. Every -// block and every top-level key inside it is REQUIRED: the binary carries no -// compiled defaults, so the adopted snapshot is exactly what the files say. +// block and every top-level key inside it is REQUIRED, but dedupe.retention +// (missing means "0", forever): the binary carries no other compiled default, +// so the adopted snapshot is exactly what the files say. // Defaults live in the seed directory (see Seed) that `wavehouse // bootstrap` writes. The fields are pointers only so Validate can tell // "absent" from the zero value and report it by path. diff --git a/internal/settings/store.go b/internal/settings/store.go index d707451d..c2609276 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -16,9 +16,9 @@ import ( // accessors below each resolve from a single snapshot load, so a reload lands // between lookups, never inside one. // -// There are no compiled defaults here on purpose: every key is required by -// Validate, so the snapshot is exactly what the files said when they were -// adopted. Defaults live in the seed directory (Seed / WriteSeed). +// There are no compiled defaults here on purpose, but one: every key but +// dedupe.retention (missing means "0", forever) is required by Validate, so +// the snapshot is exactly what the files said when they were adopted. Defaults live in the seed directory (Seed / WriteSeed). type Store struct { // tenant is the id the Registry created the store for; the zero value // only for a Store built outside a Registry (tests). From bea9e70150400f08f30d4b6bd29dd78f92dc3fce Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:34:42 -0400 Subject: [PATCH 039/108] docs(settings): the last every-key-required comment; rewrap Store's Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/settings/store.go | 3 ++- internal/settings/validate_test.go | 5 +++-- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/internal/settings/store.go b/internal/settings/store.go index c2609276..a0354641 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -18,7 +18,8 @@ import ( // // There are no compiled defaults here on purpose, but one: every key but // dedupe.retention (missing means "0", forever) is required by Validate, so -// the snapshot is exactly what the files said when they were adopted. Defaults live in the seed directory (Seed / WriteSeed). +// the snapshot is exactly what the files said when they were adopted. +// Defaults live in the seed directory (Seed / WriteSeed). type Store struct { // tenant is the id the Registry created the store for; the zero value // only for a Store built outside a Registry (tests). diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index 7af9214d..96e1ff60 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -33,8 +33,9 @@ func validFiles() map[string]string { // configJSON returns the seed config.json with patch merged over it, one // level deep (a patched block's keys replace the seed's, the rest of the -// block is kept). Every key is required, so tests that care about one key -// build a complete document from the seed rather than repeating all of them. +// block is kept). Every key but dedupe.retention is required, so tests that +// care about one key build a complete document from the seed rather than +// repeating all of them. func configJSON(patch string) string { seed, err := Seed() if err != nil { From ce14795fe88195f060ceb63f51f6e27816903892 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:04:51 -0400 Subject: [PATCH 040/108] refactor(keyenc): one escaping for subject and cache key tokens NATS subject tokens (internal/mq) and cache namespace tokens (query.SafeEncodeToken) each carried a copy of the same encoder. Both now call internal/keyenc: Escape keeps [A-Za-z0-9_] and writes every other byte as %XX, Unescape decodes as url.PathUnescape did, and Join/Split join escaped fields with a separator the escaping never emits. Output is byte-identical, pinned by golden subject tests and against v0.1.0's encoder for every byte value. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 3 +- CHANGELOG.md | 2 + docs/src/content/docs/architecture.md | 7 +- internal/keyenc/keyenc.go | 123 ++++++++++++++++++++++ internal/keyenc/keyenc_test.go | 142 ++++++++++++++++++++++++++ internal/mq/mq.go | 5 +- internal/mq/subject.go | 30 +----- internal/mq/subject_test.go | 73 ++++++------- internal/query/ident.go | 26 +---- 9 files changed, 315 insertions(+), 96 deletions(-) create mode 100644 internal/keyenc/keyenc.go create mode 100644 internal/keyenc/keyenc_test.go diff --git a/AGENTS.md b/AGENTS.md index 3d780d92..aca624a0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -26,7 +26,7 @@ One binary: - **`cmd/wavehouse/`** — Standalone mode (all-in-one with embedded NATS, optional Pebble dedup): argv dispatch, the logger, `config.Load`, and the signal context; everything else is `internal/app` -Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): +Nineteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it @@ -38,6 +38,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_]` and writes every other byte as `%XX`, `Unescape` reverses it leniently (as `url.PathUnescape` does), `Join`/`Split` join escaped fields with a separator the escaping never emits. NATS subject tokens and the cache's namespace tokens use it; its output is pinned byte for byte, since v0.1.0's subjects carry it - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index 70ce51e1..91416480 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,6 +32,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed +- **One escaping for composite keys** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject}.go` (+ tests), `internal/query/ident.go`, `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too. Behaviour-preserving: every subject and cache key is byte-identical (pinned by golden tests over dotted, wildcard, whitespace, `-`, `%`, `/`, brace, NUL and non-ASCII names, and against v0.1.0's encoder for every byte value), and a subject token decodes exactly as `url.PathUnescape` decoded it. + - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fe0c94c9..b987938c 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -61,6 +61,7 @@ internal/ ├── dedupe/ Optional deduplication (Pebble) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper +├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache namespace tokens) ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) ├── observability/ OpenTelemetry pipeline (traces/metrics/logs + Prometheus exposition) ├── pipes/ Named query pipes (NamedQuery type, parameter binding, Source) @@ -146,7 +147,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. -- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. - **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. @@ -200,6 +201,10 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi - **chsql.go** — Dependency-free ClickHouse SQL helpers shared by `query/` and `policy/`, kept in their own package to break an import cycle. `QuoteIdent` is the single place every identifier — column, table, alias — becomes SQL text: always backtick-quoted and escaped, so any ClickHouse-legal name (dots, spaces, unicode, keywords) is safe. `BindUnsafe` reports whether a name contains a literal `?`, which would desync clickhouse-go's positional binder; such names are rejected fail-closed rather than silently mis-bound. +### `keyenc/` — Key Escaping + +- **keyenc.go** — The one escaping every composite key is built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. + ## Data Flows ### Ingest Path diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go new file mode 100644 index 00000000..ddda72ca --- /dev/null +++ b/internal/keyenc/keyenc.go @@ -0,0 +1,123 @@ +// Package keyenc is the one escaping every composite WaveHouse key is built +// from: NATS subject tokens, cache namespace tokens and dedupe keys. A field +// keeps ASCII letters, digits and '_' as they are and writes every other byte +// as %XX (uppercase hex), so no separator, wildcard, whitespace, brace or +// non-ASCII byte ever appears in it unescaped, and any table name ClickHouse +// accepts encodes. +// +// The output is pinned byte for byte: NATS subjects have carried it since +// v0.1.0, and queued messages outlive the binary that wrote them. +package keyenc + +import ( + "errors" + "fmt" + "strings" +) + +const upperHex = "0123456789ABCDEF" + +// kept reports whether b is written as itself. +func kept(b byte) bool { + return (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' +} + +// Escape encodes s as one field. +func Escape(s string) string { + for i := 0; i < len(s); i++ { + if !kept(s[i]) { + return string(AppendEscape(make([]byte, 0, len(s)+2*(len(s)-i)), s)) + } + } + return s +} + +// AppendEscape appends Escape(s) to dst. +func AppendEscape(dst []byte, s string) []byte { + for i := 0; i < len(s); i++ { + b := s[i] + if kept(b) { + dst = append(dst, b) + } else { + dst = append(dst, '%', upperHex[b>>4], upperHex[b&0x0F]) + } + } + return dst +} + +// ErrBadEscape is a '%' not followed by two hex digits. +var ErrBadEscape = errors.New("keyenc: malformed escape") + +// Unescape reverses Escape. It decodes %XX in either hex case and takes any +// other byte as itself — what url.PathUnescape accepts — so a field some +// other writer left partly unescaped still reads. +func Unescape(s string) (string, error) { + i := strings.IndexByte(s, '%') + if i < 0 { + return s, nil + } + out := make([]byte, 0, len(s)) + out = append(out, s[:i]...) + for ; i < len(s); i++ { + if s[i] != '%' { + out = append(out, s[i]) + continue + } + if i+2 >= len(s) { + return "", fmt.Errorf("%w in %q", ErrBadEscape, s) + } + hi, ok1 := unhex(s[i+1]) + lo, ok2 := unhex(s[i+2]) + if !ok1 || !ok2 { + return "", fmt.Errorf("%w in %q", ErrBadEscape, s) + } + out = append(out, hi<<4|lo) + i += 2 + } + return string(out), nil +} + +func unhex(c byte) (byte, bool) { + switch { + case c >= '0' && c <= '9': + return c - '0', true + case c >= 'a' && c <= 'f': + return c - 'a' + 10, true + case c >= 'A' && c <= 'F': + return c - 'A' + 10, true + } + return 0, false +} + +// Join escapes each field and joins them with sep. It panics if sep is a +// byte Escape keeps, since a key split on it could then not be told apart. +func Join(sep byte, fields ...string) string { + return string(AppendJoin(nil, sep, fields...)) +} + +// AppendJoin appends Join(sep, fields...) to dst. +func AppendJoin(dst []byte, sep byte, fields ...string) []byte { + if kept(sep) || sep == '%' { + panic(fmt.Sprintf("keyenc: %q cannot separate fields", sep)) + } + for i, f := range fields { + if i > 0 { + dst = append(dst, sep) + } + dst = AppendEscape(dst, f) + } + return dst +} + +// Split reverses Join: the fields of key, each unescaped. +func Split(key string, sep byte) ([]string, error) { + parts := strings.Split(key, string(sep)) + for i, p := range parts { + f, err := Unescape(p) + if err != nil { + return nil, err + } + parts[i] = f + } + return parts, nil +} diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go new file mode 100644 index 00000000..7c178c86 --- /dev/null +++ b/internal/keyenc/keyenc_test.go @@ -0,0 +1,142 @@ +package keyenc_test + +import ( + "bytes" + "fmt" + "net/url" + "strings" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/keyenc" +) + +// v010Escape is the encoder v0.1.0 shipped as query.SafeEncodeNATS, copied +// verbatim: the reference Escape must match byte for byte. +func v010Escape(raw string) string { + var buf bytes.Buffer + for i := 0; i < len(raw); i++ { + b := raw[i] + if (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' { + buf.WriteByte(b) + } else { + fmt.Fprintf(&buf, "%%%02X", b) + } + } + return buf.String() +} + +func TestEscape_Golden(t *testing.T) { + t.Parallel() + for raw, want := range map[string]string{ + "": "", + "my_table123": "my_table123", + "default.clicks": "default%2Eclicks", + "my table": "my%20table", + "a-b/c": "a%2Db%2Fc", + "a.*.>": "a%2E%2A%2E%3E", + "100%": "100%25", + "{acme}:x|y": "%7Bacme%7D%3Ax%7Cy", + "a\x00b": "a%00b", + "café": "caf%C3%A9", + "\xff": "%FF", + "tab\tnewline\n": "tab%09newline%0A", + "#hash": "%23hash", + "ABCxyz_0189": "ABCxyz_0189", + "evt-123": "evt%2D123", + "日本": "%E6%97%A5%E6%9C%AC", + } { + assert.Equal(t, want, keyenc.Escape(raw), "%q", raw) + assert.Equal(t, want, string(keyenc.AppendEscape([]byte("x"), raw))[1:], "%q", raw) + } +} + +// Every byte value, alone and between kept bytes, encodes as v0.1.0 did. +func TestEscape_MatchesV010EveryByte(t *testing.T) { + t.Parallel() + for b := 0; b < 256; b++ { + for _, s := range []string{string([]byte{byte(b)}), "a" + string([]byte{byte(b)}) + "Z"} { + require.Equal(t, v010Escape(s), keyenc.Escape(s), "byte %#x", b) + } + } +} + +func TestEscape_NoAllocWhenNothingToEscape(t *testing.T) { + s := "events_2026" + assert.Zero(t, testing.AllocsPerRun(100, func() { _ = keyenc.Escape(s) })) +} + +func TestUnescape(t *testing.T) { + t.Parallel() + for in, want := range map[string]string{ + "": "", + "plain": "plain", + "default%2Eclicks": "default.clicks", + "lower%2ecase": "lower.case", + "left-as-is": "left-as-is", + "%00%FF": "\x00\xff", + "a+b": "a+b", + } { + got, err := keyenc.Unescape(in) + require.NoError(t, err, "%q", in) + assert.Equal(t, want, got, "%q", in) + } + for _, bad := range []string{"%", "%2", "a%2Gb", "%%41", "x%"} { + _, err := keyenc.Unescape(bad) + require.ErrorIs(t, err, keyenc.ErrBadEscape, "%q", bad) + } +} + +func TestJoinSplit(t *testing.T) { + t.Parallel() + assert.Equal(t, "acme/clicks/evt%2D123", keyenc.Join('/', "acme", "clicks", "evt-123")) + assert.Equal(t, "a%2Fb/c", keyenc.Join('/', "a/b", "c"), "a separator inside a field is escaped") + assert.Equal(t, "a..", keyenc.Join('.', "a", "", "")) + assert.Equal(t, "p:a", string(keyenc.AppendJoin([]byte("p:"), '/', "a"))) + + for _, fields := range [][]string{{"acme", "a/b", "id"}, {"", "", ""}, {"%", "/", "%2F"}, {"x"}} { + got, err := keyenc.Split(keyenc.Join('/', fields...), '/') + require.NoError(t, err) + assert.Equal(t, fields, got) + } + _, err := keyenc.Split("a/%zz", '/') + require.ErrorIs(t, err, keyenc.ErrBadEscape) + + for _, sep := range []byte{'a', 'Z', '5', '_', '%'} { + assert.Panics(t, func() { keyenc.Join(sep, "x") }, "%q", sep) + } +} + +// Unescape accepts exactly what url.PathUnescape did, so readers moved onto +// it decode every subject they decoded before. +func FuzzUnescapeMatchesPathUnescape(f *testing.F) { + for _, s := range []string{"", "a%2Eb", "%2e", "%", "%zz", "a+b", "%E6%97%A5", "100%25"} { + f.Add(s) + } + f.Fuzz(func(t *testing.T, s string) { + want, wantErr := url.PathUnescape(s) + got, err := keyenc.Unescape(s) + if wantErr != nil { + require.Error(t, err) + return + } + require.NoError(t, err) + require.Equal(t, want, got) + }) +} + +func FuzzEscapeRoundTrip(f *testing.F) { + for _, s := range []string{"", "a.b", "\x00", "café", "%", "a/b c"} { + f.Add(s) + } + f.Fuzz(func(t *testing.T, s string) { + enc := keyenc.Escape(s) + require.Equal(t, v010Escape(s), enc) + require.False(t, strings.ContainsAny(enc, "./:{}|# *>\x00"), enc) + dec, err := keyenc.Unescape(enc) + require.NoError(t, err) + require.Equal(t, s, dec) + }) +} diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 3f1c45c1..9198b17b 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -13,6 +13,7 @@ import ( "errors" "time" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/observability" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -38,9 +39,9 @@ type Topic struct { // as its table (parseTopicKey's fallback). Callers key their own maps by the // Topic value itself. func (t Topic) key() string { - key := string(t.Tenant) + "." + encodeToken(t.Table) + key := string(t.Tenant) + "." + keyenc.Escape(t.Table) if t.Scope != "" { - key += "." + encodeToken(t.Scope) + key += "." + keyenc.Escape(t.Scope) } return key } diff --git a/internal/mq/subject.go b/internal/mq/subject.go index 489ad241..2919464a 100644 --- a/internal/mq/subject.go +++ b/internal/mq/subject.go @@ -1,11 +1,10 @@ package mq import ( - "bytes" "fmt" - "net/url" "strings" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -59,29 +58,6 @@ func streamTenant(prefix, name string) (tenant.ID, bool) { return id, err == nil } -// encodeToken converts any table or scope name into a safe, single NATS -// subject token. It preserves alphanumerics and underscores, but -// percent-encodes everything else (so '.', ' ', '*' and '>' can never split -// or wildcard a subject). -func encodeToken(raw string) string { - var buf bytes.Buffer - for i := 0; i < len(raw); i++ { - b := raw[i] - if (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' { - buf.WriteByte(b) - } else { - fmt.Fprintf(&buf, "%%%02X", b) - } - } - return buf.String() -} - -// decodeToken reverses encodeToken. url.PathUnescape handles exactly the %XX -// form encodeToken writes. -func decodeToken(safe string) (string, error) { - return url.PathUnescape(safe) -} - // subject renders a caller's topic under prefix. The tenant is checked // against its grammar here, on the way to the wire: an empty one — a caller // that never set it — must not become a subject of some other tenant's, and @@ -118,10 +94,10 @@ func parseTopicKey(tail string) Topic { switch len(parts) { case 2, 3: id, idErr := tenant.Parse(parts[0]) - table, tableErr := decodeToken(parts[1]) + table, tableErr := keyenc.Unescape(parts[1]) scope, scopeErr := "", error(nil) if len(parts) == 3 { - scope, scopeErr = decodeToken(parts[2]) + scope, scopeErr = keyenc.Unescape(parts[2]) } if idErr == nil && tableErr == nil && scopeErr == nil { return Topic{Tenant: id, Table: table, Scope: scope} diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index 67536e4e..c2a314b4 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -9,56 +9,41 @@ import ( "github.com/stretchr/testify/require" ) -func TestEncodeToken(t *testing.T) { +// The subjects are pinned byte for byte: an embedded broker holds messages +// under them across an upgrade, and the table token has been this encoding +// since v0.1.0. +func TestSubject_Golden(t *testing.T) { t.Parallel() - tests := []struct { - name string - raw string - expected string + for _, tt := range []struct { + topic Topic + want string }{ - {"safe string", "my_table123", "my_table123"}, - {"with dots", "default.clicks", "default%2Eclicks"}, - {"with spaces", "my table", "my%20table"}, - {"with dashes and slashes", "a-b/c", "a%2Db%2Fc"}, - {"wildcards cannot survive", "a.*.>", "a%2E%2A%2E%3E"}, - {"empty string", "", ""}, - {"only safe characters", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_"}, - } - - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - assert.Equal(t, tt.expected, encodeToken(tt.raw)) - }) + {Topic{Tenant: "0", Table: "events"}, "0.events"}, + {Topic{Tenant: "acme-co", Table: "default.clicks", Scope: "org_1"}, "acme-co.default%2Eclicks.org_1"}, + {Topic{Tenant: "a", Table: "a.*.>", Scope: "*"}, "a.a%2E%2A%2E%3E.%2A"}, + {Topic{Tenant: "a", Table: "my table", Scope: "tab\there"}, "a.my%20table.tab%09here"}, + {Topic{Tenant: "a", Table: "table-with-dashes", Scope: "org-1"}, "a.table%2Dwith%2Ddashes.org%2D1"}, + {Topic{Tenant: "a", Table: "100%", Scope: "a/b"}, "a.100%25.a%2Fb"}, + {Topic{Tenant: "a", Table: "{acme}:x"}, "a.%7Bacme%7D%3Ax"}, + {Topic{Tenant: "a", Table: "nul\x00", Scope: "\xff"}, "a.nul%00.%FF"}, + {Topic{Tenant: "a", Table: "caf\u00e9", Scope: "\u65e5"}, "a.caf%C3%A9.%E6%97%A5"}, + {Topic{Tenant: "a", Table: ""}, "a."}, + {Topic{Tenant: "a", Table: "t", Scope: ""}, "a.t"}, + } { + assert.Equal(t, tt.want, tt.topic.key(), "%+v", tt.topic) + for _, prefix := range []string{ingestPrefix, dlqPrefix} { + subj, err := subject(prefix, tt.topic) + require.NoError(t, err) + assert.Equal(t, prefix+tt.want, subj) + } } } -func TestDecodeToken(t *testing.T) { +// A token another writer left partly unescaped, or escaped in lowercase, +// still reads as it always did. +func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() - tests := []struct { - name string - safe string - expected string - wantErr bool - }{ - {"safe string", "my_table123", "my_table123", false}, - {"encoded dots", "default%2Eclicks", "default.clicks", false}, - {"encoded spaces", "my%20table", "my table", false}, - {"invalid percent encoding", "default%2Gclicks", "", true}, // %2G is not valid hex - } - - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - got, err := decodeToken(tt.safe) - if tt.wantErr { - assert.Error(t, err) - } else { - require.NoError(t, err) - assert.Equal(t, tt.expected, got) - } - }) - } + assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) } func TestSubject_RoundTripsEveryTopic(t *testing.T) { diff --git a/internal/query/ident.go b/internal/query/ident.go index 320fe097..35f032da 100644 --- a/internal/query/ident.go +++ b/internal/query/ident.go @@ -1,24 +1,8 @@ package query -import ( - "bytes" - "fmt" -) +import "github.com/Wave-RF/WaveHouse/internal/keyenc" -// SafeEncodeToken converts any table or scope name into a single dot-free -// token, for composing the cache's dotted namespace keys. It preserves -// alphanumerics and underscores, but percent-encodes everything else. -func SafeEncodeToken(raw string) string { - var buf bytes.Buffer - for i := 0; i < len(raw); i++ { - b := raw[i] - // Pass through safe characters: a-z, A-Z, 0-9, and _ (underscore) - if (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' { - buf.WriteByte(b) - } else { - // Hex encode everything else (e.g., '.' becomes '%2E', ' ' becomes '%20') - fmt.Fprintf(&buf, "%%%02X", b) - } - } - return buf.String() -} +// SafeEncodeToken renders a table or scope name as one dot-free token of the +// cache's namespace keys: keyenc's escaping, the same bytes a NATS subject +// carries for the name. +func SafeEncodeToken(raw string) string { return keyenc.Escape(raw) } From 2f119735a25dcfa4025c465c6b2aeb5f5a4e56e6 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:07:48 -0400 Subject: [PATCH 041/108] refactor(dedupe): key ids as readable escaped text The dedupe key is now /
/, the table and id escaped by internal/keyenc (the NATS subject-token escaping), which never writes '/' or '#'. An id whose escaped form exceeds 1,024 bytes is stored as '#' plus its SHA-256 in hex. Keys are ASCII with no NUL, so they read as-is in a console, are valid DynamoDB Strings, and never meet the tenant-NUL-id keys written before #222. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/dedupe/dedupetest/dedupetest.go | 11 +- internal/dedupe/embedded.go | 4 +- internal/dedupe/key.go | 58 +++++----- internal/dedupe/key_layout_test.go | 108 ++++++++++++------- 7 files changed, 112 insertions(+), 75 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9a0164ee..f126ea04 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so a table name is no longer refused for the bytes it holds: one holding a NUL byte dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key is readable text, `/
/` (for example `acme/clicks/evt%2D123`), with the table and id escaped by `internal/keyenc`, the escaping NATS subject tokens already use, so a table name is no longer refused for the bytes it holds: one holding a NUL byte or a `/` dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index d609fb20..b9dedb71 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -124,7 +124,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `dedupe/` — Deduplication (Optional) - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. -- **key.go** — the key layout every backend stores: a version byte, the tenant's length and the tenant, the table's length and the table, then the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The lengths (uvarints) rather than a separator mark where each field ends, so a table name may hold any byte — NUL included — and neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. +- **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 905ba816..d849b323 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,7 +186,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit or `_` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index a9934001..67710a11 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -310,14 +310,18 @@ var cases = []struct { {"any table name is its own keyspace", func(t *testing.T, s *suite) { d := s.store(t, "acme") // seen[i] and fresh[i] differ only in where table ends and id - // begins; the first two pairs would share one key under a - // NUL-separated layout. + // begins, or in a byte an escaped key could confuse with its + // escape: each pair would share one key under a layout that + // separated the fields without escaping them. seen := []dedupe.Key{ {Table: "a", ID: "b\x00c"}, {Table: "a\x00", ID: "b"}, {Table: "", ID: "\x01a"}, {Table: "\xff\xfe", ID: "e1"}, {Table: "tab\tle \n", ID: "e1"}, + {Table: "a/b", ID: "c"}, + {Table: "a%2Fb", ID: "c"}, + {Table: "t", ID: "%23x"}, } fresh := []dedupe.Key{ {Table: "a\x00b", ID: "c"}, @@ -325,6 +329,9 @@ var cases = []struct { {Table: "\x01", ID: "a"}, {Table: "\xff", ID: "\xfee1"}, {Table: "tab\tle", ID: " \ne1"}, + {Table: "a", ID: "b/c"}, + {Table: "a/b", ID: "c%"}, + {Table: "t", ID: "#x"}, } require.NoError(t, d.Commit(t.Context(), reserve(t, d, long, seen...), 0)) for _, c := range reserve(t, d, long, fresh...) { diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index b3f015f0..5174f998 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -125,8 +125,8 @@ type tenantStore struct { } // Committed values are committedMark ‖ expiry (big-endian UnixNano, 0 = -// never). Version-0 values were a bare 8-byte timestamp under version-0 -// keys, which no version-1 key reads. +// never). Values written before #222 were a bare 8-byte timestamp, so one +// under a key that happens to equal a current one reads as absent. const ( committedMark = 2 valueLen = 9 diff --git a/internal/dedupe/key.go b/internal/dedupe/key.go index c1cfc97e..449b7418 100644 --- a/internal/dedupe/key.go +++ b/internal/dedupe/key.go @@ -2,60 +2,58 @@ package dedupe import ( "crypto/sha256" - "encoding/binary" + "encoding/hex" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/tenant" ) -// The key layout every backend stores, byte for byte: +// The key every backend stores is text: // -// keyVersion ‖ uvarint(len(tenant)) ‖ tenant ‖ uvarint(len(table)) ‖ table ‖ id +// /
/ acme/clicks/evt%2D123 +// /
/# an id too long to store verbatim // -// Each field before the id carries its length, so a table name may hold any -// byte — NUL included — and no two (tenant, table, id) triples share a key. -// The id is last, so it needs no length and may hold anything too. A tenant -// id never starts with keyVersion (tenant.Parse admits letters, digits, '_' -// and '-'), so the version-0 keys before #222 (tenant ‖ 0x00 ‖ id) never meet -// these. +// The table and id are escaped by internal/keyenc, which never writes '/' or +// '#', and a tenant id holds neither (tenant.Parse), so the fields split back +// apart, a table name may hold any byte, and no two (tenant, table, id) +// triples share a key. A key is ASCII, so it is a valid DynamoDB String, and +// holds no NUL, so it never meets a tenant ‖ NUL ‖ id key written before #222. const ( - keyVersion byte = 0x01 - // hashedID leads an id stored as its SHA-256 rather than verbatim. Ids - // that start with it are hashed too, so a verbatim id never reads as a - // hashed one. - hashedID byte = 0xFF - // MaxIDBytes is the longest id stored verbatim: a DynamoDB partition key - // holds at most 2,048 bytes, and the tenant and table share them. + keySep = '/' + hashedMark = '#' + // MaxIDBytes is the longest escaped id stored verbatim: a DynamoDB + // partition key holds at most 2,048 bytes, and the tenant and table share + // them. MaxIDBytes = 1024 ) // KeyPrefix is the part of every key that names tenant id, so a backend // computes it once per tenant store. func KeyPrefix(id tenant.ID) []byte { - p := make([]byte, 0, len(id)+1+binary.MaxVarintLen64) - p = append(p, keyVersion) - return appendField(p, string(id)) + return append([]byte(id), keySep) } // Hashed reports whether k's id is stored as its SHA-256 rather than -// verbatim. +// verbatim: whether its escaped form is longer than MaxIDBytes. func (k Key) Hashed() bool { - return len(k.ID) > MaxIDBytes || (k.ID != "" && k.ID[0] == hashedID) + switch { + case len(k.ID) > MaxIDBytes: + return true + case 3*len(k.ID) <= MaxIDBytes: // escaping at most triples a byte + return false + } + return len(keyenc.Escape(k.ID)) > MaxIDBytes } // AppendKey appends k's stored form, under the tenant prefix from KeyPrefix, // to dst. func AppendKey(dst, prefix []byte, k Key) []byte { dst = append(dst, prefix...) - dst = appendField(dst, k.Table) + dst = keyenc.AppendEscape(dst, k.Table) + dst = append(dst, keySep) if k.Hashed() { sum := sha256.Sum256([]byte(k.ID)) - dst = append(dst, hashedID) - return append(dst, sum[:]...) + return hex.AppendEncode(append(dst, hashedMark), sum[:]) } - return append(dst, k.ID...) -} - -func appendField(dst []byte, s string) []byte { - dst = binary.AppendUvarint(dst, uint64(len(s))) - return append(dst, s...) + return keyenc.AppendEscape(dst, k.ID) } diff --git a/internal/dedupe/key_layout_test.go b/internal/dedupe/key_layout_test.go index 224382ba..4f5a6e5e 100644 --- a/internal/dedupe/key_layout_test.go +++ b/internal/dedupe/key_layout_test.go @@ -2,71 +2,98 @@ package dedupe_test import ( "crypto/sha256" - "encoding/binary" + "encoding/hex" "strings" "testing" + "unicode/utf8" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/tenant" ) +func key(tn tenant.ID, k dedupe.Key) string { + return string(dedupe.AppendKey(nil, dedupe.KeyPrefix(tn), k)) +} + // The layout is pinned byte for byte: DynamoDB items and Pebble keys outlive // the binary that wrote them. func TestKeyLayout(t *testing.T) { t.Parallel() - got := dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "a\x00b", ID: "e1"}) - assert.Equal(t, []byte("\x01\x04acme\x03a\x00be1"), got) + for _, tt := range []struct { + tn tenant.ID + k dedupe.Key + want string + }{ + {"acme", dedupe.Key{Table: "clicks", ID: "evt-123"}, "acme/clicks/evt%2D123"}, + {"acme-co", dedupe.Key{Table: "db.t", ID: "a/b"}, "acme-co/db%2Et/a%2Fb"}, + {"0", dedupe.Key{Table: "a\x00b", ID: "e1"}, "0/a%00b/e1"}, + {"0", dedupe.Key{Table: "", ID: ""}, "0//"}, + {"0", dedupe.Key{Table: "t", ID: "#%café"}, "0/t/%23%25caf%C3%A9"}, + } { + assert.Equal(t, tt.want, key(tt.tn, tt.k), "%+v", tt.k) + } long := strings.Repeat("x", dedupe.MaxIDBytes+1) sum := sha256.Sum256([]byte(long)) - got = dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "t", ID: long}) - assert.Equal(t, append([]byte("\x01\x04acme\x01t\xff"), sum[:]...), got) + assert.Equal(t, "acme/t/#"+hex.EncodeToString(sum[:]), key("acme", dedupe.Key{Table: "t", ID: long})) +} - table := strings.Repeat("t", 200) - got = dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: table, ID: "e1"}) - assert.Equal(t, []byte("\x01\x04acme\xc8\x01"+table+"e1"), got, "a length past 127 takes two bytes") +// The limit is on the escaped id: 1,024 bytes that escape to more are hashed, +// and an id escaping to exactly 1,024 is not. +func TestKeyLayout_HashesOnTheEscapedLength(t *testing.T) { + t.Parallel() + fits := strings.Repeat("x", dedupe.MaxIDBytes) + assert.False(t, dedupe.Key{ID: fits}.Hashed()) + assert.Equal(t, "a/t/"+fits, key("a", dedupe.Key{Table: "t", ID: fits})) + + escapedFits := strings.Repeat("-", dedupe.MaxIDBytes/3) + "x" // 1,023 + 1 bytes escaped + assert.False(t, dedupe.Key{ID: escapedFits}.Hashed()) + escapedOver := strings.Repeat("-", dedupe.MaxIDBytes/3+1) // 1,026 bytes escaped + assert.True(t, dedupe.Key{ID: escapedOver}.Hashed()) + assert.True(t, strings.HasPrefix(key("a", dedupe.Key{Table: "t", ID: escapedOver}), "a/t/#")) } -// decodeKey inverts AppendKey. That it exists — every key parses back to the -// one triple that wrote it — is what makes the layout collision-free. -func decodeKey(t *testing.T, b []byte) (tn, table, idPart string) { +// decodeKey inverts AppendKey. That it exists — every key splits back into +// the one triple that wrote it — is what makes the layout collision-free. +func decodeKey(t *testing.T, s string) (tn, table, idPart string) { t.Helper() - require.NotEmpty(t, b) - require.Equal(t, byte(0x01), b[0]) - b = b[1:] - field := func() string { - n, w := binary.Uvarint(b) - require.Positive(t, w) - b = b[w:] - require.LessOrEqual(t, n, uint64(len(b))) - s := string(b[:n]) - b = b[n:] - return s + require.True(t, utf8.ValidString(s), "a key is a valid DynamoDB String") + require.NotContains(t, s, "\x00") + parts := strings.Split(s, "/") + require.Len(t, parts, 3, s) + _, err := tenant.Parse(parts[0]) + require.NoError(t, err) + table, err = keyenc.Unescape(parts[1]) + require.NoError(t, err) + if strings.HasPrefix(parts[2], "#") { + return parts[0], table, parts[2] } - tn = field() - table = field() - return tn, table, string(b) + id, err := keyenc.Unescape(parts[2]) + require.NoError(t, err) + return parts[0], table, id } -// Triples a separator could confuse — NUL or the version byte in the table, -// in the id, at either end, or moved across the table/id boundary — each get -// a key of their own, and parse back to themselves. +// Triples a separator could confuse — the separator, the escape and hash +// marks, NUL, or what an escape looks like, in the table or the id, at either +// end, or moved across the table/id boundary — each get a key of their own +// and parse back to themselves. func TestKeyLayout_NoCollisions(t *testing.T) { t.Parallel() - tenants := []tenant.ID{"a", "ab", "a_b", "acme"} - pieces := []string{"", "\x00", "\x01", "\xff", "a", "b", "a\x00", "\x00b", "a\x00b", "\x01\x04acme", "\x03a"} + tenants := []tenant.ID{"a", "ab", "a_b", "a-b", "acme"} + pieces := []string{"", "/", "%", "#", "\x00", "\xff", "a", "b", "a/", "/b", "a/b", "%2F", "a%2Fb", "#a", "acme/a", "é"} seen := map[string]string{} check := func(tn tenant.ID, k dedupe.Key) { - key := dedupe.AppendKey(nil, dedupe.KeyPrefix(tn), k) - who := string(tn) + " / " + k.Table + " / " + k.ID - if prev, ok := seen[string(key)]; ok && prev != who { - t.Fatalf("%q and %q share the key %q", prev, who, key) + s := key(tn, k) + who := string(tn) + " | " + k.Table + " | " + k.ID + if prev, ok := seen[s]; ok && prev != who { + t.Fatalf("%q and %q share the key %q", prev, who, s) } - seen[string(key)] = who - gotTenant, gotTable, idPart := decodeKey(t, key) + seen[s] = who + gotTenant, gotTable, idPart := decodeKey(t, s) assert.Equal(t, string(tn), gotTenant) assert.Equal(t, k.Table, gotTable) if !k.Hashed() { @@ -87,17 +114,22 @@ func TestKeyLayout_NoCollisions(t *testing.T) { grow = func(prefix string) { all = append(all, prefix) if len(prefix) < 3 { - for _, c := range []string{"\x00", "\x01", "\x04", "\xff", "a"} { + for _, c := range []string{"/", "%", "#", "2", "F", "\x00", "a"} { grow(prefix + c) } } } grow("") - for _, tn := range tenants { + for _, tn := range tenants[:2] { for _, table := range all { for _, id := range all { check(tn, dedupe.Key{Table: table, ID: id}) } } } + // A hashed id never reads as a verbatim one, even one spelling the hash. + long := strings.Repeat("x", dedupe.MaxIDBytes+1) + sum := sha256.Sum256([]byte(long)) + check("a", dedupe.Key{Table: "t", ID: long}) + check("a", dedupe.Key{Table: "t", ID: "#" + hex.EncodeToString(sum[:])}) } From 51db11caa49196195e89a32a3a92fa3d94d08f53 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:10:02 -0400 Subject: [PATCH 042/108] docs(keyenc): name only the keys built from it today Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/keyenc/keyenc.go | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index ddda72ca..807cf435 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,9 +1,9 @@ -// Package keyenc is the one escaping every composite WaveHouse key is built -// from: NATS subject tokens, cache namespace tokens and dedupe keys. A field -// keeps ASCII letters, digits and '_' as they are and writes every other byte -// as %XX (uppercase hex), so no separator, wildcard, whitespace, brace or -// non-ASCII byte ever appears in it unescaped, and any table name ClickHouse -// accepts encodes. +// Package keyenc is the one escaping composite WaveHouse keys are built +// from: NATS subject tokens and cache namespace tokens. A field keeps ASCII +// letters, digits and '_' as they are and writes every other byte as %XX +// (uppercase hex), so no separator, wildcard, whitespace, brace or non-ASCII +// byte ever appears in it unescaped, and any table name ClickHouse accepts +// encodes. // // The output is pinned byte for byte: NATS subjects have carried it since // v0.1.0, and queued messages outlive the binary that wrote them. From 27f760893f52012183723fb9d1b80ad7db4bc2f2 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:11:38 -0400 Subject: [PATCH 043/108] docs(architecture): keyenc is not yet every composite key's escaping Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/architecture.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index b987938c..43b66bd2 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -203,7 +203,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping every composite key is built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. ## Data Flows From cfdc06cd8fae71fc8167f25d8cff63195bb47c52 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:13:12 -0400 Subject: [PATCH 044/108] feat(dedupe): key the DynamoDB table by a String pk The dedupe key is now ASCII text, so the table's partition key is a String: it reads as-is in the console and in get-item output, with the same 2,048-byte limit. Check() now requires pk to be S, CreateTable declares it S, and the Deployment table and Terraform example follow. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 4 ++-- internal/dedupe/dynamodb.go | 20 ++++++++++---------- internal/dedupe/dynamodb_test.go | 16 ++++++++-------- tests/integration/dedupe_dynamodb_test.go | 12 ++++++------ 6 files changed, 28 insertions(+), 28 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a0062f44..e79d07b9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6a63f13d..8995a4bd 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index c7e88c2c..a612e3ad 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -433,7 +433,7 @@ What the backend requires of the table: | Attribute | Type | Role | |---|---|---| -| `pk` | Binary | Partition key, and the only key: tenant, table and id. No sort key. | +| `pk` | String | Partition key, and the only key: tenant, table and id as readable text, for example `acme/clicks/evt%2D123` (the table and id escaped the way NATS subject tokens are). No sort key. | | `st` | Number | `1` = pending claim, `2` = committed. | | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | @@ -451,7 +451,7 @@ resource "aws_dynamodb_table" "wavehouse_dedupe" { attribute { name = "pk" - type = "B" + type = "S" } ttl { diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index eb1c8383..bb39c060 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -187,7 +187,7 @@ func (d *Dynamo) Tenant(id tenant.ID) *Managed { } // Check verifies the table exists with the key schema this backend writes: -// pk, binary, as the only key. TTL not enabled on ex is logged, not refused: +// pk, a string, as the only key. TTL not enabled on ex is logged, not refused: // expiry never depends on it, only storage does. func (d *Dynamo) Check(ctx context.Context) error { ctx, cancel := context.WithTimeout(ctx, 10*d.cfg.Timeout) @@ -201,8 +201,8 @@ func (d *Dynamo) Check(ctx context.Context) error { return fmt.Errorf("dedupe: dynamodb table %s: key schema must be %s (HASH) alone", d.cfg.Table, attrKey) } for _, a := range t.AttributeDefinitions { - if aws.ToString(a.AttributeName) == attrKey && a.AttributeType != types.ScalarAttributeTypeB { - return fmt.Errorf("dedupe: dynamodb table %s: %s must be binary (B), is %s", d.cfg.Table, attrKey, a.AttributeType) + if aws.ToString(a.AttributeName) == attrKey && a.AttributeType != types.ScalarAttributeTypeS { + return fmt.Errorf("dedupe: dynamodb table %s: %s must be a string (S), is %s", d.cfg.Table, attrKey, a.AttributeType) } } ttl, err := d.api.DescribeTimeToLive(ctx, &dynamodb.DescribeTimeToLiveInput{TableName: &d.cfg.Table}) @@ -229,7 +229,7 @@ func (d *Dynamo) CreateTable(ctx context.Context) error { _, err := d.api.CreateTable(ctx, &dynamodb.CreateTableInput{ TableName: &d.cfg.Table, BillingMode: types.BillingModePayPerRequest, - AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String(attrKey), AttributeType: types.ScalarAttributeTypeB}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String(attrKey), AttributeType: types.ScalarAttributeTypeS}}, KeySchema: []types.KeySchemaElement{{AttributeName: aws.String(attrKey), KeyType: types.KeyTypeHash}}, }) var inUse *types.ResourceInUseException @@ -332,7 +332,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati func (s *dynamoStore) reserve(ctx context.Context, k Key, token, nowSec string, exp int64) (Status, error) { item := map[string]types.AttributeValue{ - attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, k)}, + attrKey: &types.AttributeValueMemberS{Value: string(AppendKey(nil, s.prefix, k))}, attrState: &types.AttributeValueMemberN{Value: statePending}, attrExpiry: &types.AttributeValueMemberN{Value: strconv.FormatInt(exp, 10)}, attrToken: &types.AttributeValueMemberB{Value: []byte(token)}, @@ -375,13 +375,13 @@ func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time seen := make(map[string]bool, len(claims)) writes := make([]types.WriteRequest, 0, len(claims)) for _, c := range claims { - pk := AppendKey(nil, s.prefix, c.Key) - if seen[string(pk)] { + pk := string(AppendKey(nil, s.prefix, c.Key)) + if seen[pk] { continue } - seen[string(pk)] = true + seen[pk] = true item := map[string]types.AttributeValue{ - attrKey: &types.AttributeValueMemberB{Value: pk}, + attrKey: &types.AttributeValueMemberS{Value: pk}, attrState: &types.AttributeValueMemberN{Value: stateCommitted}, attrToken: &types.AttributeValueMemberB{Value: []byte(c.Token)}, } @@ -442,7 +442,7 @@ func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { err := s.d.call(ctx, "delete_item", func(ctx context.Context) error { _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ TableName: &s.d.cfg.Table, - Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, + Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberS{Value: string(AppendKey(nil, s.prefix, c.Key))}}, ConditionExpression: aws.String(condRelease), ExpressionAttributeValues: map[string]types.AttributeValue{ ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index dc74d9d1..d2f5f740 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -137,7 +137,7 @@ func TestClassify(t *testing.T) { func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { t.Parallel() _, m := openFake(t, &fakeDynamo{put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { - id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Item[attrKey].(*types.AttributeValueMemberS).Value switch id[len(id)-1] { case 'd': return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: stateCommitted}}} @@ -162,7 +162,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { var released []string _, m := openFake(t, &fakeDynamo{ put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { - id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Item[attrKey].(*types.AttributeValueMemberS).Value mu.Lock() putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) mu.Unlock() @@ -175,7 +175,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { return &dynamodb.PutItemOutput{}, nil }, del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { - id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Key[attrKey].(*types.AttributeValueMemberS).Value mu.Lock() defer mu.Unlock() assert.Equal(t, putTokens[id], string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value), "released by the token it was put with") @@ -210,7 +210,7 @@ func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { defer mu.Unlock() var left []types.WriteRequest for i, r := range reqs { - pk := string(r.PutRequest.Item[attrKey].(*types.AttributeValueMemberB).Value) + pk := r.PutRequest.Item[attrKey].(*types.AttributeValueMemberS).Value assert.Equal(t, stateCommitted, r.PutRequest.Item[attrState].(*types.AttributeValueMemberN).Value) assert.Contains(t, r.PutRequest.Item, attrExpiry) if i == len(reqs)-1 && !heldBack[pk] && len(reqs) > 1 { @@ -252,7 +252,7 @@ func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { t.Parallel() _, m := openFake(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { assert.Equal(t, condRelease, aws.ToString(in.ConditionExpression)) - id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Key[attrKey].(*types.AttributeValueMemberS).Value if id[len(id)-1] == 'x' { return nil, &types.ResourceNotFoundException{} } @@ -320,7 +320,7 @@ func TestDynamo_Check(t *testing.T) { t.Parallel() good := &dynamodb.DescribeTableOutput{Table: &types.TableDescription{ KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, - AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeB}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeS}}, }} ttlOn := &dynamodb.DescribeTimeToLiveOutput{TimeToLiveDescription: &types.TimeToLiveDescription{ AttributeName: aws.String("ex"), TimeToLiveStatus: types.TimeToLiveStatusEnabled, @@ -375,8 +375,8 @@ func TestExpiresAt(t *testing.T) { } func idOf(av types.AttributeValue) string { - b := av.(*types.AttributeValueMemberB).Value - return string(b[len(b)-2:]) + s := av.(*types.AttributeValueMemberS).Value + return s[len(s)-2:] } func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go index 192734a9..79bea18a 100644 --- a/tests/integration/dedupe_dynamodb_test.go +++ b/tests/integration/dedupe_dynamodb_test.go @@ -222,11 +222,11 @@ func TestDedupeDynamo_Check(t *testing.T) { _, err := raw.CreateTable(t.Context(), &dynamodb.CreateTableInput{ TableName: aws.String(table), BillingMode: types.BillingModePayPerRequest, - AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeS}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeB}}, KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, }) require.NoError(t, err) - assert.ErrorContains(t, dynamoClient(t, table, dedupe.DynamoConfig{}).Check(t.Context()), "must be binary") + assert.ErrorContains(t, dynamoClient(t, table, dedupe.DynamoConfig{}).Check(t.Context()), "must be a string") fresh := newDynamoTable() d := dynamoClient(t, fresh, dedupe.DynamoConfig{}) @@ -248,13 +248,13 @@ func TestDedupeDynamo_Expiry(t *testing.T) { require.NoError(t, d.CreateTable(t.Context())) m := d.Tenant("acme") require.NoError(t, m.Apply(true)) - pk := func(id string) []byte { - return dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "events", ID: id}) + pk := func(id string) string { + return string(dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "events", ID: id})) } item := func(id string) map[string]types.AttributeValue { out, err := raw.GetItem(t.Context(), &dynamodb.GetItemInput{ TableName: aws.String(table), ConsistentRead: aws.Bool(true), - Key: map[string]types.AttributeValue{"pk": &types.AttributeValueMemberB{Value: pk(id)}}, + Key: map[string]types.AttributeValue{"pk": &types.AttributeValueMemberS{Value: pk(id)}}, }) require.NoError(t, err) return out.Item @@ -280,7 +280,7 @@ func TestDedupeDynamo_Expiry(t *testing.T) { // TTL deletes lazily; an item whose ex has passed is absent all the same. for _, st := range []string{"1", "2"} { _, err = raw.PutItem(t.Context(), &dynamodb.PutItemInput{TableName: aws.String(table), Item: map[string]types.AttributeValue{ - "pk": &types.AttributeValueMemberB{Value: pk("stale-" + st)}, + "pk": &types.AttributeValueMemberS{Value: pk("stale-" + st)}, "st": &types.AttributeValueMemberN{Value: st}, "ex": &types.AttributeValueMemberN{Value: strconv.FormatInt(time.Now().Add(-time.Minute).Unix(), 10)}, "tk": &types.AttributeValueMemberB{Value: []byte("old")}, From 78e24aa994f52a8f53b924c0304507205c638800 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:16:48 -0400 Subject: [PATCH 045/108] docs(keyenc): name the dedupe keys among its users Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 ++-- internal/keyenc/keyenc.go | 10 +++++----- 4 files changed, 9 insertions(+), 9 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 428aac54..2cde92bf 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_]` and writes every other byte as `%XX`, `Unescape` reverses it leniently (as `url.PathUnescape` does), `Join`/`Split` join escaped fields with a separator the escaping never emits. NATS subject tokens and the cache's namespace tokens use it; its output is pinned byte for byte, since v0.1.0's subjects carry it +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_]` and writes every other byte as `%XX`, `Unescape` reverses it leniently (as `url.PathUnescape` does), `Join`/`Split` join escaped fields with a separator the escaping never emits. NATS subject tokens, the cache's namespace tokens and the dedupe keys (`/
/`) use it; its output is pinned byte for byte, since v0.1.0's subjects carry it - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index f126ea04..69a67d76 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **One escaping for composite keys** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject}.go` (+ tests), `internal/query/ident.go`, `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too. Behaviour-preserving: every subject and cache key is byte-identical (pinned by golden tests over dotted, wildcard, whitespace, `-`, `%`, `/`, brace, NUL and non-ASCII names, and against v0.1.0's encoder for every byte value), and a subject token decodes exactly as `url.PathUnescape` decoded it. +- **One escaping for composite keys** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject}.go` (+ tests), `internal/query/ident.go`, `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys use too. Behaviour-preserving: every subject and cache key is byte-identical (pinned by golden tests over dotted, wildcard, whitespace, `-`, `%`, `/`, brace, NUL and non-ASCII names, and against v0.1.0's encoder for every byte value), and a subject token decodes exactly as `url.PathUnescape` decoded it. - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 096cf510..bfa48dca 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -61,7 +61,7 @@ internal/ ├── dedupe/ Optional deduplication (Reserve/Commit/Release; Pebble) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper -├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache namespace tokens) +├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache namespace tokens, dedupe keys) ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) ├── observability/ OpenTelemetry pipeline (traces/metrics/logs + Prometheus exposition) ├── pipes/ Named query pipes (NamedQuery type, parameter binding, Source) @@ -205,7 +205,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`), the cache's namespace tokens (`query.SafeEncodeToken`) and the dedupe keys (`internal/dedupe`, `/`-separated) use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. ## Data Flows diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 807cf435..e067539a 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,9 +1,9 @@ // Package keyenc is the one escaping composite WaveHouse keys are built -// from: NATS subject tokens and cache namespace tokens. A field keeps ASCII -// letters, digits and '_' as they are and writes every other byte as %XX -// (uppercase hex), so no separator, wildcard, whitespace, brace or non-ASCII -// byte ever appears in it unescaped, and any table name ClickHouse accepts -// encodes. +// from: NATS subject tokens, cache namespace tokens and dedupe keys. A field +// keeps ASCII letters, digits and '_' as they are and writes every other byte +// as %XX (uppercase hex), so no separator, wildcard, whitespace, brace or +// non-ASCII byte ever appears in it unescaped, and any table name ClickHouse +// accepts encodes. // // The output is pinned byte for byte: NATS subjects have carried it since // v0.1.0, and queued messages outlive the binary that wrote them. From 35989e488de11cd93758900adee34c7e9a018de1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:23:53 -0400 Subject: [PATCH 046/108] test(dedupe): pair / with its escape in the keyspace case; review fixes The conformance case now pairs a table holding '/' with one holding its escape, and asserts every seen key is claimed on first sight, so a backend whose keys collide fails it. CHANGELOG states what the key gives rather than a refusal no release had. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- internal/dedupe/dedupetest/dedupetest.go | 10 ++++++---- internal/settings/validate_test.go | 3 ++- 3 files changed, 9 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 69a67d76..3ca87e2b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key is readable text, `/
/` (for example `acme/clicks/evt%2D123`), with the table and id escaped by `internal/keyenc`, the escaping NATS subject tokens already use, so a table name is no longer refused for the bytes it holds: one holding a NUL byte or a `/` dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key is readable text, `/
/` (for example `acme/clicks/evt%2D123`), with the table and id escaped by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index 67710a11..821296c2 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -320,7 +320,6 @@ var cases = []struct { {Table: "\xff\xfe", ID: "e1"}, {Table: "tab\tle \n", ID: "e1"}, {Table: "a/b", ID: "c"}, - {Table: "a%2Fb", ID: "c"}, {Table: "t", ID: "%23x"}, } fresh := []dedupe.Key{ @@ -329,11 +328,14 @@ var cases = []struct { {Table: "\x01", ID: "a"}, {Table: "\xff", ID: "\xfee1"}, {Table: "tab\tle", ID: " \ne1"}, - {Table: "a", ID: "b/c"}, - {Table: "a/b", ID: "c%"}, + {Table: "a%2Fb", ID: "c"}, {Table: "t", ID: "#x"}, } - require.NoError(t, d.Commit(t.Context(), reserve(t, d, long, seen...), 0)) + first := reserve(t, d, long, seen...) + for _, c := range first { + require.Equal(t, dedupe.Claimed, c.Status, "%q shares a key with another seen key", c.Key) + } + require.NoError(t, d.Commit(t.Context(), first, 0)) for _, c := range reserve(t, d, long, fresh...) { assert.Equal(t, dedupe.Claimed, c.Status, "%q", c.Key) } diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index d9dd3479..bdb8fa6e 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -344,7 +344,8 @@ func TestValidate_ContentRules(t *testing.T) { } // A table name with odd bytes — NUL included — is any other table name to -// the override maps: dedupe keys carry the table's length, not a separator. +// the override maps: dedupe keys escape the table (internal/keyenc), so any +// bytes are just another table name. func TestValidate_OverrideTableNamesAnyBytes(t *testing.T) { t.Parallel() files := validFiles() From b299d9365fd5f3611a6d5c3703ee4334c5a809aa Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:28:39 -0400 Subject: [PATCH 047/108] test(dedupe): keep the table/id boundary pair across '/' Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dedupetest/dedupetest.go | 2 ++ 1 file changed, 2 insertions(+) diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index 821296c2..8994a795 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -320,6 +320,7 @@ var cases = []struct { {Table: "\xff\xfe", ID: "e1"}, {Table: "tab\tle \n", ID: "e1"}, {Table: "a/b", ID: "c"}, + {Table: "a/b", ID: "d"}, {Table: "t", ID: "%23x"}, } fresh := []dedupe.Key{ @@ -329,6 +330,7 @@ var cases = []struct { {Table: "\xff", ID: "\xfee1"}, {Table: "tab\tle", ID: " \ne1"}, {Table: "a%2Fb", ID: "c"}, + {Table: "a", ID: "b/d"}, {Table: "t", ID: "#x"}, } first := reserve(t, d, long, seen...) From 5d77211de17e38810cbbf3335a5fd99012fedbaa Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:37:28 -0400 Subject: [PATCH 048/108] test(dedupe): a legacy value under a current key reads as absent The version-0 read test planted only a tenant-NUL-id key, which the text layout never looks up, so it passed whatever Reserve made of a legacy value. It now also plants a bare legacy id that spells a current key and checks it reads as absent and is overwritten by the commit. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/deployment.md | 2 +- internal/dedupe/embedded_test.go | 20 ++++++++++++++------ internal/dedupe/sweep.go | 2 +- 3 files changed, 16 insertions(+), 8 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index ca23f0fe..6c103c0d 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -419,7 +419,7 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres ## Upgrading across the dedupe key change -The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys are never read, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys never count as seen, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. The same release adds an optional **`dedupe.retention`** key. No upgrade step is needed: a `config.json` without it keeps every id forever, as before. See [Deduplication](/settings-directory#deduplication) for a finite one. diff --git a/internal/dedupe/embedded_test.go b/internal/dedupe/embedded_test.go index 5015318f..8dfd2973 100644 --- a/internal/dedupe/embedded_test.go +++ b/internal/dedupe/embedded_test.go @@ -140,17 +140,25 @@ func TestEmbedded_OpenFailure(t *testing.T) { assert.True(t, e.Open()) } -// Keys from before the table joined the key (#222) are never read: an id -// seen then is accepted once more after the upgrade, the documented cost of -// the new layout. -func TestEmbedded_VersionZeroKeysAreNotRead(t *testing.T) { +// Keys from before the table joined the key (#222) never count: an id seen +// then is accepted once more after the upgrade, the documented cost of the +// new layout. A tenant ‖ NUL ‖ id key is never looked up; a bare v0.1.0 id +// that spells a current key is, and its 8-byte value reads as absent. +func TestEmbedded_VersionZeroKeysDoNotCount(t *testing.T) { t.Parallel() e := NewEmbedded(t.TempDir()) m := switchedOn(t, e, "acme") require.NoError(t, e.db.Set([]byte("acme\x00e1"), make([]byte, 8), pebble.Sync)) - dup, err := mark(context.Background(), m, "e1") + stale := AppendKey(nil, KeyPrefix("acme"), Key{Table: "events", ID: "e2"}) + require.NoError(t, e.db.Set(stale, make([]byte, 8), pebble.Sync)) + for _, id := range []string{"e1", "e2"} { + dup, err := mark(context.Background(), m, id) + require.NoError(t, err) + assert.False(t, dup, id) + } + dup, err := mark(context.Background(), m, "e2") require.NoError(t, err) - assert.False(t, dup) + assert.True(t, dup, "the commit overwrote the stale value") } // A claim nobody commits, releases or reserves again leaves memory at the diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index 444ec178..70018684 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -75,7 +75,7 @@ func (e *Embedded) startSweep(db *pebble.DB) (stop func()) { } // sweep makes one pass over the whole instance, deleting keys whose -// retention has ended and version-0 keys, which nothing reads: those from +// retention has ended and version-0 keys, which never count: those from // before ids were keyed by table (tenant ‖ 0x00 ‖ id, or the bare id before // that). They are told apart by value, since a bare id may be any bytes, a // current key's included: only commits are stored, and every commit has the From 7082907a576486b2948aad21b2c029d08a17e92d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:44:33 -0400 Subject: [PATCH 049/108] fix(dedupe): let a sent DynamoDB put answer before its release Reserve sent every put on the errgroup's context, so the first failure cancelled siblings already on the wire. The client gave up on them but the table could still apply one after the undo's conditional delete had found nothing, holding the key for the whole lease: the conformance case "a failed reserve leaves nothing claimed" failed about one run in ten against dynamodb-local. A sent put now runs on the caller's context; the group's context only skips the puts not yet sent. 30 consecutive conformance runs against dynamodb-local pass. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dynamodb.go | 9 +++++--- internal/dedupe/dynamodb_test.go | 35 +++++++++++++++++++------------- 2 files changed, 27 insertions(+), 17 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index bb39c060..5b444d5e 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -292,8 +292,11 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati claims := make([]Claim, len(keys)) tried := make([]Claim, len(keys)) sent := make([]bool, len(keys)) - // The first failure cancels the puts not yet sent: the Reserve fails - // either way, and a throttled table should not take the rest. + // The first failure skips the puts not yet sent: the Reserve fails + // either way, and a throttled table should not take the rest. A put + // already sent runs on ctx, not gctx, so it finishes and its outcome is + // known before the undo below; cancelled mid-flight, it could land after + // its release and hold the key for the lease. g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { @@ -304,7 +307,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati return err } sent[i] = true - status, err := s.reserve(gctx, k, token, nowSec, exp) + status, err := s.reserve(ctx, k, token, nowSec, exp) if err != nil { return err } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index d2f5f740..685cb7f7 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -419,12 +419,14 @@ func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { assert.Equal(t, int64(4), deletes.Load()) } -// One throttled put in a multi-key Reserve: the unsent puts are neither sent -// nor released, and the cancelled siblings do not reset the breaker. +// One throttled put in a multi-key Reserve: the unsent puts are never sent, +// and a sibling already sent runs to its answer before the undo releases it, +// so a put cannot land after its own release. func TestDynamo_FailedMultiKeyReserve(t *testing.T) { t.Parallel() var mu sync.Mutex var put, released []string + landed := map[string]bool{} _, m := openFakeWith(t, &fakeDynamo{ put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := idOf(in.Item[attrKey]) @@ -434,26 +436,31 @@ func TestDynamo_FailedMultiKeyReserve(t *testing.T) { if id == "k0" { return nil, &types.ProvisionedThroughputExceededException{} } - <-ctx.Done() - return nil, ctx.Err() + time.Sleep(20 * time.Millisecond) + if err := ctx.Err(); err != nil { + return nil, err + } + mu.Lock() + landed[id] = true + mu.Unlock() + return &dynamodb.PutItemOutput{}, nil }, del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := idOf(in.Key[attrKey]) mu.Lock() - released = append(released, idOf(in.Key[attrKey])) - mu.Unlock() - return nil, &types.ConditionalCheckFailedException{} + defer mu.Unlock() + released = append(released, id) + if id != "k0" && !landed[id] { + t.Errorf("%s released before its put answered", id) + } + return &dynamodb.DeleteItemOutput{}, nil }, }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) ks := keys("k0", "k1", "k2", "k3", "k4", "k5", "k6", "k7") - for range breakerTrips { - _, err := m.Reserve(t.Context(), ks, time.Minute) - require.ErrorIs(t, err, ErrUnavailable) - require.NotErrorIs(t, err, errBreakerOpen) - } _, err := m.Reserve(t.Context(), ks, time.Minute) - require.ErrorIs(t, err, errBreakerOpen, "the cancelled siblings did not reset the count") + require.ErrorIs(t, err, ErrUnavailable) mu.Lock() defer mu.Unlock() assert.ElementsMatch(t, put, released, "exactly the sent puts are released") - assert.Less(t, len(put), breakerTrips*len(ks), "unsent puts were never sent") + assert.Less(t, len(put), len(ks), "unsent puts were never sent") } From 484c72e056a876aca9118f1a804d343715c36ab9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:45:08 -0400 Subject: [PATCH 050/108] docs(changelog): the boot-backends entry predates dedupe.dynamodb Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 65f18ee9..3e557140 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer's in-process backend is its default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings go in a `.` sub-block (`dedupe.dynamodb` is the first, entry above); any other sub-block is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. From 4c2c2a407f52b4386c748d0d2d50a6a4b8dd788f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:49:45 -0400 Subject: [PATCH 051/108] test(dedupe): a caller-cancelled put does not reset the breaker The guard's comment named sibling cancellation, which no longer happens; it now names the caller going away, and a test pins it (removing the guard fails it). Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dynamodb.go | 5 +++-- internal/dedupe/dynamodb_test.go | 30 ++++++++++++++++++++++++++++++ 2 files changed, 33 insertions(+), 2 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 5b444d5e..d298a8c7 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -260,8 +260,9 @@ func (d *Dynamo) call(ctx context.Context, op string, do func(context.Context) e start := time.Now() err := classify(op, do(ctx)) d.metrics.record(ctx, op, time.Since(start), err) - // A request cancelled because a sibling failed says nothing about the - // table, and must not reset the breaker's count. + // A request cancelled because its caller went away (a client + // disconnecting mid-Reserve) says nothing about the table, and must not + // reset the breaker's count. if op == opReserve && !errors.Is(err, context.Canceled) { d.breaker.record(err) } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 685cb7f7..88bc3f76 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -401,6 +401,36 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { assert.Equal(t, 2*batchWriteMax, written, "a failed chunk does not cancel the others: their records are published") } +// A put cancelled by its caller is not an answer from the table: it does +// not reset the breaker's count of throttled puts. +func TestDynamo_CallerCancelDoesNotResetBreaker(t *testing.T) { + t.Parallel() + var hang atomic.Bool + started := make(chan struct{}, 1) + _, m := openFake(t, &fakeDynamo{put: func(ctx context.Context, _ *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + if hang.Load() { + started <- struct{}{} + <-ctx.Done() + return nil, ctx.Err() + } + return nil, &types.ProvisionedThroughputExceededException{} + }}) + for range breakerTrips - 1 { + _, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + } + hang.Store(true) + ctx, cancel := context.WithCancel(t.Context()) + go func() { <-started; cancel() }() + _, err := m.Reserve(ctx, keys("a"), time.Minute) + require.ErrorIs(t, err, context.Canceled) + hang.Store(false) + _, err = m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + _, err = m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, errBreakerOpen, "the cancelled put did not reset the count") +} + func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { t.Parallel() var deletes atomic.Int64 From 2c31eff00e5ed4c507c514bf950e0b1057c1f594 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:30:42 -0400 Subject: [PATCH 052/108] refactor(keyenc): keep '-', join keys, count dead letters per table Escape now keeps '-' alongside [A-Za-z0-9_], exactly the tenant-id grammar, so a tenant id is its own escaped form and dashed names read as themselves. Unescape is url.PathUnescape, so v0.1.0's %2D still decodes. Join/AppendJoin/Split are fixed: Split converted a separator of 0x80+ as a rune, and Join of no fields could not be told from one empty field; both now refuse those inputs. Topic keys are built with AppendJoin and parsed with Split, the tenant token still verbatim. Dead-letter counts are keyed by table, every scope of a table summed under it, through one deadLetterTables function; the ?table= filter keeps all of a table's scopes. A scoped message used to count under "table.scope", which a dotted table name could share. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 4 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 5 +- internal/keyenc/keyenc.go | 77 +++++++------------ internal/keyenc/keyenc_test.go | 102 +++++++++++++++++--------- internal/mq/deadletter.go | 18 +++++ internal/mq/deadletter_test.go | 24 ++++++ internal/mq/embedded.go | 21 +----- internal/mq/mq.go | 24 +++--- internal/mq/subject.go | 37 +++++----- internal/mq/subject_test.go | 10 ++- internal/query/ident_test.go | 2 +- 12 files changed, 186 insertions(+), 140 deletions(-) create mode 100644 internal/mq/deadletter.go create mode 100644 internal/mq/deadletter_test.go diff --git a/AGENTS.md b/AGENTS.md index 30c808c9..bf8dc6a7 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,8 +38,8 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_]` and writes every other byte as `%XX`, `Unescape` reverses it leniently (as `url.PathUnescape` does), `Join`/`Split` join escaped fields with a separator the escaping never emits. NATS subject tokens and the cache's namespace tokens use it; its output is pinned byte for byte, since v0.1.0's subjects carry it -- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so v0.1.0's `%2D` still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key +- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) - **`policy/`** — Hasura-style access control, **role-first**: `TablePolicy` is `map[string]RolePermissions`, and a role's grant splits by operation into `SelectPermissions` (columns, row `filter`, aggregations, the `max_*` limits) and `InsertPermissions` (columns, `check`) — so a field only one side honors does not exist on the other. `Evaluate()` resolves ONE operation and leaves the other side **nil** (`Select *ResolvedSelect` / `Insert *ResolvedInsert`), which every accessor fails closed on — nil is "not resolved", distinct from an empty side, which is "unrestricted" (what the admin return builds). Claim templating (`{{ jwt.claim.path }}`) resolves during that call. Policies come from `Source`, a `func() *Policy` read per call (`settings.Store.Policy` in production, `Static(p)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index 96c8081c..6283af53 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **One escaping for composite keys** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject}.go` (+ tests), `internal/query/ident.go`, `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too. Behaviour-preserving: every subject and cache key is byte-identical (pinned by golden tests over dotted, wildcard, whitespace, `-`, `%`, `/`, brace, NUL and non-ASCII names, and against v0.1.0's encoder for every byte value), and a subject token decodes exactly as `url.PathUnescape` decoded it. +- **One escaping for composite keys, `-` kept; dead-letter counts per table** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject,embedded,deadletter}.go` (+ tests), `internal/query/ident.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too, and subjects are built with its `Join` (each field escaped, with a separator the escaping never writes between them). The escaping now keeps `-` as well as ASCII letters, digits and `_` — exactly the tenant-id grammar — so a table or scope such as `my-table` is `my-table` in a subject rather than `my%2Dtable`; every other byte is escaped as before (pinned by golden tests, and against v0.1.0's encoder for every other byte value). Decoding is `url.PathUnescape`, as it was, so a subject written with v0.1.0's `%2D` reads as the same topic and a dead-letter count merges both forms. One thing notices the change: a `/v1/stream` client resuming across the upgrade (`Last-Event-ID` or `since`) on a table whose name holds `-` misses that table's events queued before the upgrade, since the replay filters on the table's exact subject. `GET /v1/ops/dlq/stats` now counts every scope of a table under the table itself, and `?table=` keeps all of its scopes; a scoped message used to count under `table.scope`, a name a dotted table could share. Scope is always empty today, so the response is unchanged. - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 43b66bd2..59b2fe3f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -147,7 +147,8 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. -- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: ASCII letters, digits, `_` and `-` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **deadletter.go** — `deadLetterTables`, the per-table count `DeadLetterCounts` reports: a dead-letter stream's per-subject counts, each subject parsed back to its topic and counted under its table — every scope of a table under the table itself, so a dotted table name never shares a count with a table + scope pair — and a table filter keeps that table with all of its scopes. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. - **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. @@ -203,7 +204,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so v0.1.0's `%2D` for `-` still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. ## Data Flows diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 807cf435..59bac7e1 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,17 +1,19 @@ // Package keyenc is the one escaping composite WaveHouse keys are built // from: NATS subject tokens and cache namespace tokens. A field keeps ASCII -// letters, digits and '_' as they are and writes every other byte as %XX +// letters, digits, '_' and '-' as they are and writes every other byte as %XX // (uppercase hex), so no separator, wildcard, whitespace, brace or non-ASCII // byte ever appears in it unescaped, and any table name ClickHouse accepts -// encodes. +// encodes. The bytes it keeps are exactly a tenant id's (tenant.Parse), so a +// tenant id is its own escaped form. // -// The output is pinned byte for byte: NATS subjects have carried it since -// v0.1.0, and queued messages outlive the binary that wrote them. +// Keys built from it are stored — queued under NATS subjects, held in caches +// — so a change to what it keeps orphans them. v0.1.0 escaped '-' as %2D; +// Unescape still reads that form. package keyenc import ( - "errors" "fmt" + "net/url" "strings" ) @@ -19,7 +21,7 @@ const upperHex = "0123456789ABCDEF" // kept reports whether b is written as itself. func kept(b byte) bool { - return (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' + return (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' || b == '-' } // Escape encodes s as one field. @@ -45,60 +47,33 @@ func AppendEscape(dst []byte, s string) []byte { return dst } -// ErrBadEscape is a '%' not followed by two hex digits. -var ErrBadEscape = errors.New("keyenc: malformed escape") - -// Unescape reverses Escape. It decodes %XX in either hex case and takes any -// other byte as itself — what url.PathUnescape accepts — so a field some -// other writer left partly unescaped still reads. +// Unescape reverses Escape. It is url.PathUnescape: %XX in either hex case +// decodes, and any other byte reads as itself, so a field another writer +// left partly unescaped — v0.1.0's %2D included — still reads. func Unescape(s string) (string, error) { - i := strings.IndexByte(s, '%') - if i < 0 { - return s, nil - } - out := make([]byte, 0, len(s)) - out = append(out, s[:i]...) - for ; i < len(s); i++ { - if s[i] != '%' { - out = append(out, s[i]) - continue - } - if i+2 >= len(s) { - return "", fmt.Errorf("%w in %q", ErrBadEscape, s) - } - hi, ok1 := unhex(s[i+1]) - lo, ok2 := unhex(s[i+2]) - if !ok1 || !ok2 { - return "", fmt.Errorf("%w in %q", ErrBadEscape, s) - } - out = append(out, hi<<4|lo) - i += 2 - } - return string(out), nil + return url.PathUnescape(s) } -func unhex(c byte) (byte, bool) { - switch { - case c >= '0' && c <= '9': - return c - '0', true - case c >= 'a' && c <= 'f': - return c - 'a' + 10, true - case c >= 'A' && c <= 'F': - return c - 'A' + 10, true +// checkSep panics unless sep can separate escaped fields: a byte Escape never +// writes, and ASCII, so the key stays valid UTF-8. +func checkSep(sep byte) { + if kept(sep) || sep == '%' || sep >= 0x80 { + panic(fmt.Sprintf("keyenc: %q cannot separate fields", sep)) } - return 0, false } -// Join escapes each field and joins them with sep. It panics if sep is a -// byte Escape keeps, since a key split on it could then not be told apart. +// Join escapes each field and joins them with sep. It panics on no fields, +// whose key would be one empty field's, and on a separator Escape could +// write. func Join(sep byte, fields ...string) string { return string(AppendJoin(nil, sep, fields...)) } // AppendJoin appends Join(sep, fields...) to dst. func AppendJoin(dst []byte, sep byte, fields ...string) []byte { - if kept(sep) || sep == '%' { - panic(fmt.Sprintf("keyenc: %q cannot separate fields", sep)) + checkSep(sep) + if len(fields) == 0 { + panic("keyenc: Join needs at least one field") } for i, f := range fields { if i > 0 { @@ -109,9 +84,11 @@ func AppendJoin(dst []byte, sep byte, fields ...string) []byte { return dst } -// Split reverses Join: the fields of key, each unescaped. +// Split reverses Join: the fields of key, each unescaped. It panics on a +// separator Join would refuse. func Split(key string, sep byte) ([]string, error) { - parts := strings.Split(key, string(sep)) + checkSep(sep) + parts := strings.Split(key, string([]byte{sep})) for i, p := range parts { f, err := Unescape(p) if err != nil { diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go index 7c178c86..5b1630d6 100644 --- a/internal/keyenc/keyenc_test.go +++ b/internal/keyenc/keyenc_test.go @@ -3,7 +3,6 @@ package keyenc_test import ( "bytes" "fmt" - "net/url" "strings" "testing" @@ -11,10 +10,11 @@ import ( "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/keyenc" + "github.com/Wave-RF/WaveHouse/internal/tenant" ) // v010Escape is the encoder v0.1.0 shipped as query.SafeEncodeNATS, copied -// verbatim: the reference Escape must match byte for byte. +// verbatim. Escape differs from it only in keeping '-'. func v010Escape(raw string) string { var buf bytes.Buffer for i := 0; i < len(raw); i++ { @@ -35,7 +35,7 @@ func TestEscape_Golden(t *testing.T) { "my_table123": "my_table123", "default.clicks": "default%2Eclicks", "my table": "my%20table", - "a-b/c": "a%2Db%2Fc", + "a-b/c": "a-b%2Fc", "a.*.>": "a%2E%2A%2E%3E", "100%": "100%25", "{acme}:x|y": "%7Bacme%7D%3Ax%7Cy", @@ -45,7 +45,7 @@ func TestEscape_Golden(t *testing.T) { "tab\tnewline\n": "tab%09newline%0A", "#hash": "%23hash", "ABCxyz_0189": "ABCxyz_0189", - "evt-123": "evt%2D123", + "evt-123": "evt-123", "日本": "%E6%97%A5%E6%9C%AC", } { assert.Equal(t, want, keyenc.Escape(raw), "%q", raw) @@ -53,21 +53,52 @@ func TestEscape_Golden(t *testing.T) { } } -// Every byte value, alone and between kept bytes, encodes as v0.1.0 did. -func TestEscape_MatchesV010EveryByte(t *testing.T) { +// Every byte value, alone and between kept bytes, encodes as v0.1.0 did, +// but for '-'. +func TestEscape_MatchesV010ButDash(t *testing.T) { t.Parallel() - for b := 0; b < 256; b++ { + for b := range 256 { for _, s := range []string{string([]byte{byte(b)}), "a" + string([]byte{byte(b)}) + "Z"} { - require.Equal(t, v010Escape(s), keyenc.Escape(s), "byte %#x", b) + want := v010Escape(s) + if b == '-' { + want = s + } + require.Equal(t, want, keyenc.Escape(s), "byte %#x", b) } } } +// A tenant id is its own escaped form: the kept bytes are its grammar. +func TestEscape_KeepsExactlyTheTenantGrammar(t *testing.T) { + t.Parallel() + for b := range 256 { + s := string([]byte{byte(b)}) + _, err := tenant.Parse(s) + assert.Equal(t, err == nil, keyenc.Escape(s) == s, "byte %#x", b) + } +} + func TestEscape_NoAllocWhenNothingToEscape(t *testing.T) { - s := "events_2026" + s := "events_2026-09" assert.Zero(t, testing.AllocsPerRun(100, func() { _ = keyenc.Escape(s) })) } +// Distinct names never share an escaped form, even names that look escaped: +// '%' is itself escaped. +func TestEscape_LookalikesStayDistinct(t *testing.T) { + t.Parallel() + names := []string{"b-c", "b%2Dc", "b%2dc", "b.c", "b%2Ec"} + seen := map[string]string{} + for _, n := range names { + e := keyenc.Escape(n) + require.NotContains(t, seen, e, "%q and %q", seen[e], n) + seen[e] = n + back, err := keyenc.Unescape(e) + require.NoError(t, err) + assert.Equal(t, n, back) + } +} + func TestUnescape(t *testing.T) { t.Parallel() for in, want := range map[string]string{ @@ -75,7 +106,7 @@ func TestUnescape(t *testing.T) { "plain": "plain", "default%2Eclicks": "default.clicks", "lower%2ecase": "lower.case", - "left-as-is": "left-as-is", + "evt%2D123": "evt-123", // v0.1.0's form "%00%FF": "\x00\xff", "a+b": "a+b", } { @@ -85,55 +116,58 @@ func TestUnescape(t *testing.T) { } for _, bad := range []string{"%", "%2", "a%2Gb", "%%41", "x%"} { _, err := keyenc.Unescape(bad) - require.ErrorIs(t, err, keyenc.ErrBadEscape, "%q", bad) + require.Error(t, err, "%q", bad) } } func TestJoinSplit(t *testing.T) { t.Parallel() - assert.Equal(t, "acme/clicks/evt%2D123", keyenc.Join('/', "acme", "clicks", "evt-123")) + assert.Equal(t, "acme/clicks/evt-123", keyenc.Join('/', "acme", "clicks", "evt-123")) assert.Equal(t, "a%2Fb/c", keyenc.Join('/', "a/b", "c"), "a separator inside a field is escaped") assert.Equal(t, "a..", keyenc.Join('.', "a", "", "")) + assert.Equal(t, "", keyenc.Join('.', ""), "one empty field") assert.Equal(t, "p:a", string(keyenc.AppendJoin([]byte("p:"), '/', "a"))) - for _, fields := range [][]string{{"acme", "a/b", "id"}, {"", "", ""}, {"%", "/", "%2F"}, {"x"}} { + for _, fields := range [][]string{{"acme", "a/b", "id"}, {"", "", ""}, {""}, {"%", "/", "%2F"}, {"x"}} { got, err := keyenc.Split(keyenc.Join('/', fields...), '/') require.NoError(t, err) assert.Equal(t, fields, got) } _, err := keyenc.Split("a/%zz", '/') - require.ErrorIs(t, err, keyenc.ErrBadEscape) + require.Error(t, err) +} - for _, sep := range []byte{'a', 'Z', '5', '_', '%'} { - assert.Panics(t, func() { keyenc.Join(sep, "x") }, "%q", sep) - } +func TestJoin_RefusesZeroFields(t *testing.T) { + t.Parallel() + assert.Panics(t, func() { keyenc.Join('/') }) + assert.Panics(t, func() { keyenc.AppendJoin(nil, '/') }) } -// Unescape accepts exactly what url.PathUnescape did, so readers moved onto -// it decode every subject they decoded before. -func FuzzUnescapeMatchesPathUnescape(f *testing.F) { - for _, s := range []string{"", "a%2Eb", "%2e", "%", "%zz", "a+b", "%E6%97%A5", "100%25"} { - f.Add(s) - } - f.Fuzz(func(t *testing.T, s string) { - want, wantErr := url.PathUnescape(s) - got, err := keyenc.Unescape(s) - if wantErr != nil { - require.Error(t, err) - return +// A separator Escape could write, or one outside ASCII, is refused by Join +// and Split alike; every other byte separates. +func TestSeparators(t *testing.T) { + t.Parallel() + for b := range 256 { + sep := byte(b) + if keyenc.Escape(string([]byte{sep})) == string([]byte{sep}) || sep == '%' || sep >= 0x80 { + assert.Panics(t, func() { keyenc.Join(sep, "x") }, "%#x", b) + assert.Panics(t, func() { _, _ = keyenc.Split("x", sep) }, "%#x", b) + continue } - require.NoError(t, err) - require.Equal(t, want, got) - }) + fields := []string{"a", string([]byte{sep}), "b" + string([]byte{sep}) + "c"} + got, err := keyenc.Split(keyenc.Join(sep, fields...), sep) + require.NoError(t, err, "%#x", b) + assert.Equal(t, fields, got, "%#x", b) + } } func FuzzEscapeRoundTrip(f *testing.F) { - for _, s := range []string{"", "a.b", "\x00", "café", "%", "a/b c"} { + for _, s := range []string{"", "a.b", "\x00", "café", "%", "a/b c", "b-c", "b%2Dc"} { f.Add(s) } f.Fuzz(func(t *testing.T, s string) { enc := keyenc.Escape(s) - require.Equal(t, v010Escape(s), enc) + require.Equal(t, strings.ReplaceAll(v010Escape(s), "%2D", "-"), enc) require.False(t, strings.ContainsAny(enc, "./:{}|# *>\x00"), enc) dec, err := keyenc.Unescape(enc) require.NoError(t, err) diff --git a/internal/mq/deadletter.go b/internal/mq/deadletter.go new file mode 100644 index 00000000..e6cc98e1 --- /dev/null +++ b/internal/mq/deadletter.go @@ -0,0 +1,18 @@ +package mq + +// deadLetterTables counts a dead-letter stream's parked messages per table, +// from its per-subject counts under prefix. Every scope of a table counts +// under the table itself, so no table + scope pair can share a count with a +// dotted table name. A non-empty table keeps that table alone, all of its +// scopes included. +func deadLetterTables(subjects map[string]uint64, prefix, table string) map[string]uint64 { + tables := make(map[string]uint64, len(subjects)) + for subj, n := range subjects { + t := parseTopicKey(topicKey(prefix, subj)) + if table != "" && t.Table != table { + continue + } + tables[t.Table] += n + } + return tables +} diff --git a/internal/mq/deadletter_test.go b/internal/mq/deadletter_test.go new file mode 100644 index 00000000..aaed64fe --- /dev/null +++ b/internal/mq/deadletter_test.go @@ -0,0 +1,24 @@ +package mq + +import ( + "testing" + + "github.com/stretchr/testify/assert" +) + +func TestDeadLetterTables(t *testing.T) { + t.Parallel() + subjects := map[string]uint64{ + "dlq.0.a%2Eb": 1, // table "a.b" + "dlq.0.a.b": 2, // table "a", scope "b" + "dlq.0.a": 4, // table "a", unscoped + "dlq.0.my-t": 8, + "dlq.0.my%2Dt.org-1": 16, // v0.1.0's escaping of '-', scoped + "dlq.0.clicks.org%2E": 32, + } + assert.Equal(t, map[string]uint64{"a.b": 1, "a": 6, "my-t": 24, "clicks": 32}, + deadLetterTables(subjects, dlqPrefix, ""), "a dotted table never shares a count with a table + scope") + assert.Equal(t, map[string]uint64{"a": 6}, deadLetterTables(subjects, dlqPrefix, "a"), "the filter keeps every scope of its table") + assert.Equal(t, map[string]uint64{"a.b": 1}, deadLetterTables(subjects, dlqPrefix, "a.b")) + assert.Empty(t, deadLetterTables(subjects, dlqPrefix, "never_failed")) +} diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 98971492..f314840d 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -1003,9 +1003,9 @@ func (e *EmbeddedNATS) PurgeAcked(ctx context.Context, consumer string, olderTha } // DeadLetterCounts reads tenant id's dead-letter stream's per-subject counts -// and keys them by table. The table filter matches that table's unscoped -// subject, so it is applied to the parsed topic rather than as a subject -// filter; a scoped topic counts under "table.scope". +// and keys them by table (deadLetterTables). The table filter matches every +// scope of that table, so it is applied to the parsed topic rather than as a +// subject filter. func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) { if _, err := tenant.Parse(string(id)); err != nil { return DeadLetterCounts{}, fmt.Errorf("tenant: %w", err) @@ -1023,20 +1023,7 @@ func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, id tenant.ID, table return DeadLetterCounts{}, fmt.Errorf("dlq stream info: %w", err) } - counts := DeadLetterCounts{Tables: make(map[string]uint64, len(state.Subjects)), Total: state.Msgs} - for subj, n := range state.Subjects { - t := parseTopicKey(topicKey(dlqPrefix, subj)) - if table != "" && (t.Table != table || t.Scope != "") { - continue - } - name := t.Table - if t.Scope != "" { - // TODO(#235): break scopes out rather than fold them into the name. - name += "." + t.Scope - } - counts.Tables[name] += n - } - return counts, nil + return DeadLetterCounts{Tables: deadLetterTables(state.Subjects, dlqPrefix, table), Total: state.Msgs}, nil } // ReplaySince creates an ephemeral consumer on topic's ingest subject, in its diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 9198b17b..57eae51c 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -34,16 +34,16 @@ type Topic struct { // key is the injective string form of the topic that a subject's tail // carries: the tenant first, verbatim — its grammar makes it one token — then -// the table and scope as encoded tokens. A topic without a tenant has no -// subject, and its key parses back to a topic of no tenant with the whole key -// as its table (parseTopicKey's fallback). Callers key their own maps by the -// Topic value itself. +// the table and scope joined as escaped tokens (keyenc.AppendJoin). A topic +// without a tenant has no subject, and its key parses back to a topic of no +// tenant with the whole key as its table (parseTopicKey's fallback). Callers +// key their own maps by the Topic value itself. func (t Topic) key() string { - key := string(t.Tenant) + "." + keyenc.Escape(t.Table) - if t.Scope != "" { - key += "." + keyenc.Escape(t.Scope) + key := append([]byte(t.Tenant), '.') + if t.Scope == "" { + return string(keyenc.AppendJoin(key, '.', t.Table)) } - return key + return string(keyenc.AppendJoin(key, '.', t.Table, t.Scope)) } // Message represents a message received from the queue. @@ -247,8 +247,8 @@ type DeadLetterer interface { // DeadLetterCounts is what is parked on one tenant's dead-letter queue. type DeadLetterCounts struct { // Tables maps table name → parked messages, for the tables asked about. - // Scope is not broken out yet (it is inert until #235): a message parked - // under a scoped topic counts under "table.scope", not under its table. + // Every scope of a table counts under the table; scope is not broken out + // yet (it is inert until #235). Tables map[string]uint64 // Total is every parked message of the tenant, whatever the filter. Total uint64 @@ -263,8 +263,8 @@ var ErrNoDeadLetterQueue = errors.New("dead-letter queue not found") type DeadLetterStats interface { // DeadLetterCounts counts tenant id's parked messages per table — a // tenant served, rejected, or removed alike, for as long as its queue is - // kept. A non-empty table narrows Tables to that one (its unscoped - // messages). + // kept. A non-empty table narrows Tables to that one (all of its + // scopes). DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) } diff --git a/internal/mq/subject.go b/internal/mq/subject.go index 2919464a..2ee41ea3 100644 --- a/internal/mq/subject.go +++ b/internal/mq/subject.go @@ -84,24 +84,27 @@ func keyTenant(key string) (tenant.ID, bool) { return id, err == nil } -// parseTopicKey recovers the Topic from a subject tail. Three tokens are -// tenant, table and scope; two are tenant and table. A tail this package -// could not have written — one token, more than three, a token that does not -// decode, a tenant outside the grammar — cannot be split reliably, so the -// whole of it becomes the table of no tenant rather than being dropped. +// parseTopicKey recovers the Topic from a subject tail: the tenant token, +// read verbatim as keyTenant reads it, then one or two escaped tokens, table +// and scope (keyenc.Split). A tail this package could not have written — one +// token, more than three, a token that does not decode, a tenant outside the +// grammar — cannot be split reliably, so the whole of it becomes the table of +// no tenant rather than being dropped. func parseTopicKey(tail string) Topic { - parts := strings.Split(tail, ".") - switch len(parts) { - case 2, 3: - id, idErr := tenant.Parse(parts[0]) - table, tableErr := keyenc.Unescape(parts[1]) - scope, scopeErr := "", error(nil) - if len(parts) == 3 { - scope, scopeErr = keyenc.Unescape(parts[2]) - } - if idErr == nil && tableErr == nil && scopeErr == nil { - return Topic{Tenant: id, Table: table, Scope: scope} - } + first, rest, ok := strings.Cut(tail, ".") + if !ok { + return Topic{Table: tail} + } + id, idErr := tenant.Parse(first) + fields, fieldsErr := keyenc.Split(rest, '.') + if idErr != nil || fieldsErr != nil { + return Topic{Table: tail} + } + switch len(fields) { + case 1: + return Topic{Tenant: id, Table: fields[0]} + case 2: + return Topic{Tenant: id, Table: fields[0], Scope: fields[1]} } return Topic{Table: tail} } diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index c2a314b4..22d3bf9d 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -10,8 +10,7 @@ import ( ) // The subjects are pinned byte for byte: an embedded broker holds messages -// under them across an upgrade, and the table token has been this encoding -// since v0.1.0. +// under them across an upgrade. func TestSubject_Golden(t *testing.T) { t.Parallel() for _, tt := range []struct { @@ -22,7 +21,7 @@ func TestSubject_Golden(t *testing.T) { {Topic{Tenant: "acme-co", Table: "default.clicks", Scope: "org_1"}, "acme-co.default%2Eclicks.org_1"}, {Topic{Tenant: "a", Table: "a.*.>", Scope: "*"}, "a.a%2E%2A%2E%3E.%2A"}, {Topic{Tenant: "a", Table: "my table", Scope: "tab\there"}, "a.my%20table.tab%09here"}, - {Topic{Tenant: "a", Table: "table-with-dashes", Scope: "org-1"}, "a.table%2Dwith%2Ddashes.org%2D1"}, + {Topic{Tenant: "a", Table: "table-with-dashes", Scope: "org-1"}, "a.table-with-dashes.org-1"}, {Topic{Tenant: "a", Table: "100%", Scope: "a/b"}, "a.100%25.a%2Fb"}, {Topic{Tenant: "a", Table: "{acme}:x"}, "a.%7Bacme%7D%3Ax"}, {Topic{Tenant: "a", Table: "nul\x00", Scope: "\xff"}, "a.nul%00.%FF"}, @@ -40,10 +39,12 @@ func TestSubject_Golden(t *testing.T) { } // A token another writer left partly unescaped, or escaped in lowercase, -// still reads as it always did. +// still reads as it always did — and so does v0.1.0's %2D for '-', so a +// message queued before '-' was kept reads as the same topic. func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) + assert.Equal(t, parseTopicKey("a.table-with-dashes.org-1"), parseTopicKey("a.table%2Dwith%2Ddashes.org%2D1")) } func TestSubject_RoundTripsEveryTopic(t *testing.T) { @@ -121,6 +122,7 @@ func TestParseTopicKey_ForeignTailKeepsItself(t *testing.T) { "a.b.c.d", // more tokens than any topic renders "0.bad%2Gtoken", // a token that does not decode "a%2Eb.events", // a tenant outside the grammar + "a%2Db.events", // a tenant token is read verbatim, never decoded ".events", // a topic whose tenant was never set "events", // one token: no tenant leads it "bad%2G", // one token that does not decode diff --git a/internal/query/ident_test.go b/internal/query/ident_test.go index 572b005e..ac4fd64a 100644 --- a/internal/query/ident_test.go +++ b/internal/query/ident_test.go @@ -16,7 +16,7 @@ func TestEncodeTable(t *testing.T) { {"safe string", "my_table123", "my_table123"}, {"with dots", "default.clicks", "default%2Eclicks"}, {"with spaces", "my table", "my%20table"}, - {"with dashes and slashes", "a-b/c", "a%2Db%2Fc"}, + {"with dashes and slashes", "a-b/c", "a-b%2Fc"}, {"empty string", "", ""}, {"only safe characters", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_"}, } From 999e92c5186831fced3ff000e00c4b0ab34b3c60 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:37:09 -0400 Subject: [PATCH 053/108] docs(keyenc): pin the %2D compatibility to builds since #612 A v0.1.0 queue is deleted at boot, so the lenient %2D decoding and the stream-resume caveat only concern queues an unreleased build since #612 wrote; say so in the changelog and the comments. List keyenc in the development.md tree and AGENTS.md's file structure, and correct the stale "L2" cache line beside it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 3 ++- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/development.md | 3 ++- internal/keyenc/keyenc.go | 6 +++--- internal/keyenc/keyenc_test.go | 2 +- internal/mq/deadletter_test.go | 2 +- internal/mq/subject_test.go | 4 ++-- 8 files changed, 13 insertions(+), 11 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index bf8dc6a7..59236113 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so v0.1.0's `%2D` still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) @@ -435,6 +435,7 @@ internal/config/ → Configuration structs + loader internal/dedupe/ → Optional deduplication (interface + embedded/distributed) internal/discovery/ → ClickHouse schema introspection + ingest validation internal/ingest/ → Batch buffer with DLQ + Active Sweeper (NATS message lifecycle) +internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache namespace tokens) internal/mq/ → MQ boundary (the only NATS/JetStream importer: owned message/consumer/stream types + embedded server) internal/observability/ → OpenTelemetry pipeline (traces/metrics/logs providers, Prometheus exporter, slog fan-out, message-header trace propagation) internal/pipes/ → Named query pipes (types, parameter binding, Source) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6283af53..7861daaa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **One escaping for composite keys, `-` kept; dead-letter counts per table** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject,embedded,deadletter}.go` (+ tests), `internal/query/ident.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too, and subjects are built with its `Join` (each field escaped, with a separator the escaping never writes between them). The escaping now keeps `-` as well as ASCII letters, digits and `_` — exactly the tenant-id grammar — so a table or scope such as `my-table` is `my-table` in a subject rather than `my%2Dtable`; every other byte is escaped as before (pinned by golden tests, and against v0.1.0's encoder for every other byte value). Decoding is `url.PathUnescape`, as it was, so a subject written with v0.1.0's `%2D` reads as the same topic and a dead-letter count merges both forms. One thing notices the change: a `/v1/stream` client resuming across the upgrade (`Last-Event-ID` or `since`) on a table whose name holds `-` misses that table's events queued before the upgrade, since the replay filters on the table's exact subject. `GET /v1/ops/dlq/stats` now counts every scope of a table under the table itself, and `?table=` keeps all of its scopes; a scoped message used to count under `table.scope`, a name a dotted table could share. Scope is always empty today, so the response is unchanged. +- **One escaping for composite keys, `-` kept; dead-letter counts per table** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject,embedded,deadletter}.go` (+ tests), `internal/query/ident.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{architecture,development}.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too, and subjects are built with its `Join` (each field escaped, with a separator the escaping never writes between them). The escaping now keeps `-` as well as ASCII letters, digits and `_` — exactly the tenant-id grammar — so a table or scope such as `my-table` is `my-table` in a subject rather than `my%2Dtable`; every other byte is escaped as before (pinned by golden tests, and against v0.1.0's encoder for every other byte value). Upgrading from v0.1.0 notices nothing further, since its queue is deleted at boot (below). A queue an unreleased build since [#612](https://github.com/Wave-RF/WaveHouse/pull/612) wrote still reads, because decoding is `url.PathUnescape` as it was: `%2D` decodes to `-`, and a dead-letter count merges both forms. Only a `/v1/stream` client resuming across such an upgrade (`Last-Event-ID` or `since`) on a table whose name holds `-` misses that table's events queued before it, since the replay filters on the table's exact subject. `GET /v1/ops/dlq/stats` now counts every scope of a table under the table itself, and `?table=` keeps all of its scopes; a scoped message used to count under `table.scope`, a name a dotted table could share. Scope is always empty today, so the response is unchanged. - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 59b2fe3f..7f0aa379 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -204,7 +204,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so v0.1.0's `%2D` for `-` still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so a `%2D` for `-` that an earlier build wrote still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. ## Data Flows diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 01f82e73..16b65a74 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -454,13 +454,14 @@ WaveHouse/ │ ├── api/ # HTTP handlers, router, middleware │ ├── app/ # Process wiring (build every component, run under one errgroup, release in reverse) │ ├── auth/ # JWT/JWKS authentication middleware -│ ├── cache/ # L1 (Ristretto) + L2 caching +│ ├── cache/ # Query cache: Ristretto L1 + the tenant-led version index │ ├── chconn/ # ClickHouse pools, one per connection tuple (reconciled on settings reload) │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration │ ├── dedupe/ # Optional deduplication (Pebble) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper +│ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens) │ ├── mq/ # MQ boundary: the only NATS/JetStream importer │ ├── observability/ # OpenTelemetry pipeline (traces/metrics/logs + Prometheus) │ ├── pipes/ # Named query pipes (types + parameter binding) diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 59bac7e1..0e3434bc 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -7,8 +7,8 @@ // tenant id is its own escaped form. // // Keys built from it are stored — queued under NATS subjects, held in caches -// — so a change to what it keeps orphans them. v0.1.0 escaped '-' as %2D; -// Unescape still reads that form. +// — so a change to what it keeps orphans them. Earlier builds escaped '-' as +// %2D; Unescape still reads that form. package keyenc import ( @@ -49,7 +49,7 @@ func AppendEscape(dst []byte, s string) []byte { // Unescape reverses Escape. It is url.PathUnescape: %XX in either hex case // decodes, and any other byte reads as itself, so a field another writer -// left partly unescaped — v0.1.0's %2D included — still reads. +// left partly unescaped — an earlier build's %2D included — still reads. func Unescape(s string) (string, error) { return url.PathUnescape(s) } diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go index 5b1630d6..877599e2 100644 --- a/internal/keyenc/keyenc_test.go +++ b/internal/keyenc/keyenc_test.go @@ -106,7 +106,7 @@ func TestUnescape(t *testing.T) { "plain": "plain", "default%2Eclicks": "default.clicks", "lower%2ecase": "lower.case", - "evt%2D123": "evt-123", // v0.1.0's form + "evt%2D123": "evt-123", // an earlier build's form "%00%FF": "\x00\xff", "a+b": "a+b", } { diff --git a/internal/mq/deadletter_test.go b/internal/mq/deadletter_test.go index aaed64fe..aed76094 100644 --- a/internal/mq/deadletter_test.go +++ b/internal/mq/deadletter_test.go @@ -13,7 +13,7 @@ func TestDeadLetterTables(t *testing.T) { "dlq.0.a.b": 2, // table "a", scope "b" "dlq.0.a": 4, // table "a", unscoped "dlq.0.my-t": 8, - "dlq.0.my%2Dt.org-1": 16, // v0.1.0's escaping of '-', scoped + "dlq.0.my%2Dt.org-1": 16, // an earlier build's escaping of '-', scoped "dlq.0.clicks.org%2E": 32, } assert.Equal(t, map[string]uint64{"a.b": 1, "a": 6, "my-t": 24, "clicks": 32}, diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index 22d3bf9d..28f4940b 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -39,8 +39,8 @@ func TestSubject_Golden(t *testing.T) { } // A token another writer left partly unescaped, or escaped in lowercase, -// still reads as it always did — and so does v0.1.0's %2D for '-', so a -// message queued before '-' was kept reads as the same topic. +// still reads as it always did — and so does an earlier build's %2D for '-', +// so a message it queued reads as the same topic. func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) From 28f3886bbb11d7bee4ce1c7297913ea7c2aa1b53 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:42:07 -0400 Subject: [PATCH 054/108] refactor(dedupe): build keys with keyenc.AppendJoin; keep '-' internal/dedupe/key.go now builds the stored key with keyenc.AppendJoin instead of a manual AppendEscape + separator, so the table and id can never reach the key unescaped even as fields are added. Escape keeps '-' as of the keyenc refactor, so an id like evt-123 is stored as itself rather than evt%2D123; Hashed()'s "escaping at most triples a byte" bound still holds since the kept set only grew. key_layout_test.go: the pinned example and the escaped-length fixture (which relied on '-' always escaping to three bytes) are updated for the wider kept set, and decodeKey now reads the whole key back with keyenc.Split instead of splitting and unescaping fields by hand. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/key.go | 20 ++++++++++---------- internal/dedupe/key_layout_test.go | 20 +++++++------------- 3 files changed, 18 insertions(+), 24 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3d212957..b993acf0 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -124,7 +124,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `dedupe/` — Deduplication (Optional) - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. -- **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. +- **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt-123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped and joined (`keyenc.AppendJoin`) by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. diff --git a/internal/dedupe/key.go b/internal/dedupe/key.go index 449b7418..94654725 100644 --- a/internal/dedupe/key.go +++ b/internal/dedupe/key.go @@ -10,14 +10,15 @@ import ( // The key every backend stores is text: // -// /
/ acme/clicks/evt%2D123 +// /
/ acme/clicks/evt-123 // /
/# an id too long to store verbatim // -// The table and id are escaped by internal/keyenc, which never writes '/' or -// '#', and a tenant id holds neither (tenant.Parse), so the fields split back -// apart, a table name may hold any byte, and no two (tenant, table, id) -// triples share a key. A key is ASCII, so it is a valid DynamoDB String, and -// holds no NUL, so it never meets a tenant ‖ NUL ‖ id key written before #222. +// The table and id are escaped and joined by internal/keyenc, which never +// writes '/' or '#', and a tenant id holds neither (tenant.Parse), so the +// fields split back apart, a table name may hold any byte, and no two +// (tenant, table, id) triples share a key. A key is ASCII, so it is a valid +// DynamoDB String, and holds no NUL, so it never meets a tenant ‖ NUL ‖ id +// key written before #222. const ( keySep = '/' hashedMark = '#' @@ -49,11 +50,10 @@ func (k Key) Hashed() bool { // to dst. func AppendKey(dst, prefix []byte, k Key) []byte { dst = append(dst, prefix...) - dst = keyenc.AppendEscape(dst, k.Table) - dst = append(dst, keySep) if k.Hashed() { sum := sha256.Sum256([]byte(k.ID)) - return hex.AppendEncode(append(dst, hashedMark), sum[:]) + dst = keyenc.AppendJoin(dst, keySep, k.Table) + return hex.AppendEncode(append(dst, keySep, hashedMark), sum[:]) } - return keyenc.AppendEscape(dst, k.ID) + return keyenc.AppendJoin(dst, keySep, k.Table, k.ID) } diff --git a/internal/dedupe/key_layout_test.go b/internal/dedupe/key_layout_test.go index 4f5a6e5e..d7977f5c 100644 --- a/internal/dedupe/key_layout_test.go +++ b/internal/dedupe/key_layout_test.go @@ -28,7 +28,7 @@ func TestKeyLayout(t *testing.T) { k dedupe.Key want string }{ - {"acme", dedupe.Key{Table: "clicks", ID: "evt-123"}, "acme/clicks/evt%2D123"}, + {"acme", dedupe.Key{Table: "clicks", ID: "evt-123"}, "acme/clicks/evt-123"}, {"acme-co", dedupe.Key{Table: "db.t", ID: "a/b"}, "acme-co/db%2Et/a%2Fb"}, {"0", dedupe.Key{Table: "a\x00b", ID: "e1"}, "0/a%00b/e1"}, {"0", dedupe.Key{Table: "", ID: ""}, "0//"}, @@ -50,9 +50,9 @@ func TestKeyLayout_HashesOnTheEscapedLength(t *testing.T) { assert.False(t, dedupe.Key{ID: fits}.Hashed()) assert.Equal(t, "a/t/"+fits, key("a", dedupe.Key{Table: "t", ID: fits})) - escapedFits := strings.Repeat("-", dedupe.MaxIDBytes/3) + "x" // 1,023 + 1 bytes escaped + escapedFits := strings.Repeat(".", dedupe.MaxIDBytes/3) + "x" // 1,023 + 1 bytes escaped assert.False(t, dedupe.Key{ID: escapedFits}.Hashed()) - escapedOver := strings.Repeat("-", dedupe.MaxIDBytes/3+1) // 1,026 bytes escaped + escapedOver := strings.Repeat(".", dedupe.MaxIDBytes/3+1) // 1,026 bytes escaped assert.True(t, dedupe.Key{ID: escapedOver}.Hashed()) assert.True(t, strings.HasPrefix(key("a", dedupe.Key{Table: "t", ID: escapedOver}), "a/t/#")) } @@ -63,18 +63,12 @@ func decodeKey(t *testing.T, s string) (tn, table, idPart string) { t.Helper() require.True(t, utf8.ValidString(s), "a key is a valid DynamoDB String") require.NotContains(t, s, "\x00") - parts := strings.Split(s, "/") - require.Len(t, parts, 3, s) - _, err := tenant.Parse(parts[0]) - require.NoError(t, err) - table, err = keyenc.Unescape(parts[1]) + parts, err := keyenc.Split(s, '/') require.NoError(t, err) - if strings.HasPrefix(parts[2], "#") { - return parts[0], table, parts[2] - } - id, err := keyenc.Unescape(parts[2]) + require.Len(t, parts, 3, s) + _, err = tenant.Parse(parts[0]) require.NoError(t, err) - return parts[0], table, id + return parts[0], parts[1], parts[2] } // Triples a separator could confuse — the separator, the escape and hash From 79ab364b1d70d16ff0325cffd8413c9bdde93892 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:45:58 -0400 Subject: [PATCH 055/108] test(mq): cover a byte left unescaped in the lenient-token case With '-' kept, "b-c" is what Escape writes, so the case no longer exercised a partly unescaped token; "b~c" does. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- internal/mq/subject_test.go | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index 28f4940b..6f859acd 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -43,7 +43,7 @@ func TestSubject_Golden(t *testing.T) { // so a message it queued reads as the same topic. func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() - assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) + assert.Equal(t, Topic{Tenant: "a", Table: "b~c", Scope: "d.e"}, parseTopicKey("a.b~c.d%2ee")) assert.Equal(t, parseTopicKey("a.table-with-dashes.org-1"), parseTopicKey("a.table%2Dwith%2Ddashes.org%2D1")) } From 733588876af642dbbccd6b727940cd7421a69901 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:02:24 -0400 Subject: [PATCH 056/108] docs(dedupe): '-' is kept in the hashing bound and the keyenc listings settings-directory.mdx still said every byte but a letter, digit or '_' escapes to three; '-' is kept now. List dedupe keys among keyenc's users in development.md and the package doc, which also names what an orphaned dedupe key costs. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- CHANGELOG.md | 2 +- docs/src/content/docs/development.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/keyenc/keyenc.go | 7 ++++--- 4 files changed, 7 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index cb7e3272..c0fa9a79 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 16b65a74..75442d0f 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -461,7 +461,7 @@ WaveHouse/ │ ├── dedupe/ # Optional deduplication (Pebble) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper -│ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens) +│ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens, dedupe keys) │ ├── mq/ # MQ boundary: the only NATS/JetStream importer │ ├── observability/ # OpenTelemetry pipeline (traces/metrics/logs + Prometheus) │ ├── pipes/ # Named query pipes (types + parameter binding) diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index d849b323..b485e048 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,7 +186,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit or `_` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 0fac960b..e4c5958e 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -6,9 +6,10 @@ // ClickHouse accepts encodes. The bytes it keeps are exactly a tenant id's // (tenant.Parse), so a tenant id is its own escaped form. // -// Keys built from it are stored — queued under NATS subjects, held in caches -// — so a change to what it keeps orphans them. Earlier builds escaped '-' as -// %2D; Unescape still reads that form. +// Keys built from it are stored — queued under NATS subjects, held in caches, +// kept as dedupe keys — so a change to what it keeps orphans them; an +// orphaned dedupe key lets a seen id through again. Earlier builds escaped +// '-' as %2D; Unescape still reads that form. package keyenc import ( From bfebceeb8943f2659a7f7bb1883239923ea7824f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:29:00 -0400 Subject: [PATCH 057/108] test(dedupe): pin a pre-#222 collision and the hashed marker's escaping Round 2 review: committedLive's length/mark-byte check had no test that reached it. Add two embedded_test.go cases that write a raw Pebble value under the exact key AppendKey computes for a colliding (tenant, table, id) triple: an 8-byte v0.1.0-style value (never live, and overwritten by the next commit) and a 9-byte value with the wrong mark byte (also never live). Mutating the mark-byte check to `len(val) != valueLen { return len(val) > 0 }` fails both; reverted after confirming. dedupetest's "long ids and ids that look hashed stay distinct" case used a 0xFF-prefixed id left over from the retired length-prefixed layout. Its "looks hashed" id now spells the current hashed form's own marker and hash, escaped (`#` + sha256 hex), and the first reserve's statuses are asserted Claimed/Claimed so a collision from an unescaped-marker bug answers InFlight instead of passing silently. Mutating AppendKey's hashed branch to build the marker through AppendJoin (escaping the `#` like an ordinary field) now fails the case; reverted after confirming. internal/api/ingest.go: wavehouse_dedupe_commit_failed_total renamed to wavehouse_ingest_dedupe_commit_failed_total, matching its meter-mates wavehouse_ingest_dedupe_missing_id_total and wavehouse_ingest_dedupe_disabled_total. Updated CHANGELOG.md and settings-directory.mdx; grepped the whole Wave-RF tree for the old name and found no reference outside this repo's other worktrees (their own, unmodified copies of the same pre-rename docs and coverage artifacts). Docs: AGENTS.md's file-structure listing now lists dedupe keys among keyenc's users, and the CHANGELOG's deployment link carries the upgrade section's anchor. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/api/ingest.go | 2 +- internal/dedupe/dedupetest/dedupetest.go | 13 ++++++- internal/dedupe/embedded_test.go | 39 ++++++++++++++++++++ 6 files changed, 54 insertions(+), 6 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index bf00e5e7..07a8b6c0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -435,7 +435,7 @@ internal/config/ → Configuration structs + loader internal/dedupe/ → Optional deduplication (interface + embedded/distributed) internal/discovery/ → ClickHouse schema introspection + ingest validation internal/ingest/ → Batch buffer with DLQ + Active Sweeper (NATS message lifecycle) -internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache namespace tokens) +internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache namespace tokens, dedupe keys) internal/mq/ → MQ boundary (the only NATS/JetStream importer: owned message/consumer/stream types + embedded server) internal/observability/ → OpenTelemetry pipeline (traces/metrics/logs providers, Prometheus exporter, slog fan-out, message-header trace propagation) internal/pipes/ → Named query pipes (types, parameter binding, Source) diff --git a/CHANGELOG.md b/CHANGELOG.md index c0fa9a79..556d65a9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md#upgrading-across-the-dedupe-key-change)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_ingest_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index b485e048..c0e8d724 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,7 +186,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_ingest_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index f4051516..3b4aca60 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -82,7 +82,7 @@ var dedupeMissingIDCounter, _ = otel.Meter("wavehouse-ingest").Int64Counter( // dedupeCommitFailedCounter counts records published whose id could not be // committed afterwards: a retry after the lease lapses publishes them again. var dedupeCommitFailedCounter, _ = otel.Meter("wavehouse-ingest").Int64Counter( - "wavehouse_dedupe_commit_failed_total", + "wavehouse_ingest_dedupe_commit_failed_total", metric.WithDescription("Published records whose dedupe id failed to commit afterwards (the claim lapses with its lease)"), ) diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index 8994a795..5b9221b6 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -6,6 +6,8 @@ package dedupetest import ( "context" + "crypto/sha256" + "encoding/hex" "fmt" "strings" "sync" @@ -260,8 +262,15 @@ var cases = []struct { d := s.store(t, "acme") base := strings.Repeat("x", 2*dedupe.MaxIDBytes) longA, longB := key(base+"a"), key(base+"b") - hashLike := key("\xff" + strings.Repeat("0", 32)) - require.NoError(t, d.Commit(t.Context(), reserve(t, d, long, longA, hashLike), 0)) + // hashLike spells longA's own stored hashed form, escaped: it fits + // verbatim and stays distinct only if the '#' the hashed form writes + // raw never gets escaped like an ordinary field. + sum := sha256.Sum256([]byte(base + "a")) + hashLike := key("#" + hex.EncodeToString(sum[:])) + first := reserve(t, d, long, longA, hashLike) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed}, statuses(first), + "both fresh — a collision would answer the second InFlight") + require.NoError(t, d.Commit(t.Context(), first, 0)) assert.Equal(t, []dedupe.Status{dedupe.Duplicate, dedupe.Claimed, dedupe.Duplicate}, statuses(reserve(t, d, long, longA, longB, hashLike))) }}, diff --git a/internal/dedupe/embedded_test.go b/internal/dedupe/embedded_test.go index 5015318f..71acc68f 100644 --- a/internal/dedupe/embedded_test.go +++ b/internal/dedupe/embedded_test.go @@ -153,6 +153,45 @@ func TestEmbedded_VersionZeroKeysAreNotRead(t *testing.T) { assert.False(t, dup) } +// v0.1.0 stored a bare id as the key with an 8-byte value, so a v0.1.0 id +// that happens to spell a key the current layout would also write — +// tenant "0", table "events", id "e1" join to "0/events/e1", which a v0.1.0 +// record could have used as its own id — must not read as a live duplicate: +// only a value of exactly valueLen bytes leading with committedMark is ours. +// The reserve below claims the key despite the stale value, and only the +// commit it makes turns a second reserve of the same id into a duplicate. +func TestEmbedded_PreJoinValueUnderACollidingKeyIsNotLive(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + m := switchedOn(t, e, "0") + key := AppendKey(nil, KeyPrefix("0"), Key{Table: "events", ID: "e1"}) + require.NoError(t, e.db.Set(key, make([]byte, 8), pebble.Sync)) + + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.False(t, dup, "an 8-byte v0.1.0 value is not this layout's commit") + + dup, err = mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.True(t, dup, "the mark above overwrote it with a real commit") +} + +// Same colliding key, a 9-byte value that committedMark did not write: the +// mark byte, not just the length, is what says a value is ours. +func TestEmbedded_WrongMarkByteIsNotLive(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + m := switchedOn(t, e, "0") + key := AppendKey(nil, KeyPrefix("0"), Key{Table: "events", ID: "e1"}) + val := make([]byte, valueLen) + val[0] = committedMark + 1 + require.NoError(t, e.db.Set(key, val, pebble.Sync)) + + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.False(t, dup, "a value not led by committedMark is not a live commit") +} + // A claim nobody commits, releases or reserves again leaves memory at the // next sweep past its lease, not never. func TestPendingShard_SweepDropsLapsedClaims(t *testing.T) { From f507a2e122e8fc8ec45c09d4f28890152afe56cd Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:39:08 -0400 Subject: [PATCH 058/108] test(dedupe): pin the length half of committedLive; document the lease The planted 8-byte value now leads with committedMark, so only the length check refuses it; without that check the test panics decoding a short expiry. settings-directory.mdx now says what the 30-second lease is: a second request for an id being published gets 503 with Retry-After: 30. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- docs/src/content/docs/settings-directory.mdx | 2 +- internal/dedupe/embedded_test.go | 6 +++++- 2 files changed, 6 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e8d724..60143c28 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,7 +186,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_ingest_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. While its record is being published, an id is held for a 30-second lease: another request carrying the same id meanwhile gets `503` (`a request with the same dedupe id is in flight`) with `Retry-After: 30` — see [the ingest errors](/api#post-v1ingesttabletable--ingest-data). An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_ingest_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/dedupe/embedded_test.go b/internal/dedupe/embedded_test.go index 71acc68f..9fdefc0c 100644 --- a/internal/dedupe/embedded_test.go +++ b/internal/dedupe/embedded_test.go @@ -158,6 +158,8 @@ func TestEmbedded_VersionZeroKeysAreNotRead(t *testing.T) { // tenant "0", table "events", id "e1" join to "0/events/e1", which a v0.1.0 // record could have used as its own id — must not read as a live duplicate: // only a value of exactly valueLen bytes leading with committedMark is ours. +// The planted value leads with committedMark, so only the length check can +// refuse it (and keeps a short value from being decoded as an expiry). // The reserve below claims the key despite the stale value, and only the // commit it makes turns a second reserve of the same id into a duplicate. func TestEmbedded_PreJoinValueUnderACollidingKeyIsNotLive(t *testing.T) { @@ -165,7 +167,9 @@ func TestEmbedded_PreJoinValueUnderACollidingKeyIsNotLive(t *testing.T) { e := NewEmbedded(t.TempDir()) m := switchedOn(t, e, "0") key := AppendKey(nil, KeyPrefix("0"), Key{Table: "events", ID: "e1"}) - require.NoError(t, e.db.Set(key, make([]byte, 8), pebble.Sync)) + stale := make([]byte, 8) + stale[0] = committedMark + require.NoError(t, e.db.Set(key, stale, pebble.Sync)) dup, err := mark(context.Background(), m, "e1") require.NoError(t, err) From 155138d99d2d4cb5dadedf805a1604dd0b0824a1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:46:41 -0400 Subject: [PATCH 059/108] test(dedupe): a failed release keeps the publish's error; drop "valid" keys The Deduplicator contract still said Managed hands a backend "valid" keys, left from the removed Key.Validate. MockDeduplicator.ReleaseErr had no test: a release that fails after a failed publish is only logged, the request answers with the publish's own error, and the id lapses with its lease. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- internal/api/ingest_test.go | 18 ++++++++++++++++++ internal/dedupe/dedupe.go | 4 ++-- 2 files changed, 20 insertions(+), 2 deletions(-) diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index f08d2eb0..f452eef1 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -2789,6 +2789,24 @@ func TestIngest_Dedup_FailedPublishReleasesTheID(t *testing.T) { } } +// A release that fails after a failed publish is only logged: the request +// answers with the publish's own error, and the id is left to lapse with its +// lease rather than being reported as a dedupe failure. +func TestIngest_Dedup_FailedReleaseKeepsThePublishError(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: errors.New("connection reset")} + dedup := testutil.NewMockDeduplicator() + dedup.ReleaseErr = errors.New("store unavailable") + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}))) + require.Equal(t, http.StatusInternalServerError, w.Code) + assert.NotContains(t, w.Body.String(), "dedupe", "the publish's error, not the release's") + assert.Len(t, dedup.Released, 1, "the release was attempted") + assert.True(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "left to lapse with its lease") +} + // A batch whose publish fails part-way keeps what it published: the records // before the failure are committed, the failing one is released, and a // whole-batch retry reports the first as duplicates and publishes the rest. diff --git a/internal/dedupe/dedupe.go b/internal/dedupe/dedupe.go index 08020265..ada64a5e 100644 --- a/internal/dedupe/dedupe.go +++ b/internal/dedupe/dedupe.go @@ -58,8 +58,8 @@ type Claim struct { } // Deduplicator is a tenant's store of seen ids. Callers reach every backend -// through Managed, which hands a backend distinct, valid keys, a lease > 0, -// and only Claimed claims to Commit and Release — a backend may assume all +// through Managed, which hands a backend distinct keys, a lease > 0, and +// only Claimed claims to Commit and Release — a backend may assume all // three, and Managed's callers get the behaviour below either way. // // Reserve is atomic per key: of any number of concurrent Reserves for the From 44e0a69733726b35234afd972ff63ef504c1084c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:54:24 -0400 Subject: [PATCH 060/108] test(dedupe): cover Status.String Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- internal/dedupe/dedupe_test.go | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) create mode 100644 internal/dedupe/dedupe_test.go diff --git a/internal/dedupe/dedupe_test.go b/internal/dedupe/dedupe_test.go new file mode 100644 index 00000000..a2655bb3 --- /dev/null +++ b/internal/dedupe/dedupe_test.go @@ -0,0 +1,19 @@ +package dedupe + +import ( + "testing" + + "github.com/stretchr/testify/assert" +) + +func TestStatus_String(t *testing.T) { + t.Parallel() + for s, want := range map[Status]string{ + Claimed: "claimed", + Duplicate: "duplicate", + InFlight: "in_flight", + Status(0): "unknown", + } { + assert.Equal(t, want, s.String()) + } +} From 3248e7eadaf786bf9144dc1b40bf6570564911b0 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:47:57 -0400 Subject: [PATCH 061/108] fix(dedupe): a retried DynamoDB put keeps its own claim When the SDK retries a conditional PutItem whose first attempt was applied (a 500 or a reset connection after the write), the retry fails its condition on the caller's own item. Reserve read that as another request's live claim and answered InFlight: nobody would commit or release the item, the id stayed locked for the lease, and the producer got a spurious 503. A pending item carrying the put's own token now answers Claimed, so the caller commits it, or a failed Reserve's undo releases it. Pinned by a unit test (the fake hands back the put's own item) and an integration test that forwards the first PutItem to dynamodb-local and answers it 500. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 18 +++++++-- internal/dedupe/dynamodb_test.go | 48 +++++++++++++++++++++++ tests/integration/dedupe_dynamodb_test.go | 46 ++++++++++++++++++++++ 5 files changed, 111 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a81bf89a..792cb65d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7482a7b6..9fb7dd8f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index d298a8c7..5e71db3f 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -277,8 +277,9 @@ type dynamoStore struct { // Reserve puts every key's pending item in parallel, each conditional on no // live item holding the key. A failed condition hands back the live item, -// whose state says Duplicate or InFlight without a read. On any error every -// put that may have landed is released by its token. +// whose state says Duplicate or InFlight without a read, or Claimed when its +// token is the put's own. On any error every put that may have landed is +// released by its token. func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { if len(keys) == 0 { return []Claim{}, nil @@ -358,12 +359,23 @@ func (s *dynamoStore) reserve(ctx context.Context, k Key, token, nowSec string, } return Claimed, nil } - if st, ok := held.Item[attrState].(*types.AttributeValueMemberN); ok && st.Value == stateCommitted { + st, _ := held.Item[attrState].(*types.AttributeValueMemberN) + switch { + case st != nil && st.Value == stateCommitted: return Duplicate, nil + case st != nil && st.Value == statePending && heldBy(held.Item, token): + // This put's own item: an SDK retry of an attempt that was applied + // but whose answer was lost (a 500, a connection reset). + return Claimed, nil } return InFlight, nil } +func heldBy(item map[string]types.AttributeValue, token string) bool { + tk, ok := item[attrToken].(*types.AttributeValueMemberB) + return ok && string(tk.Value) == token +} + // Commit overwrites every claim's item as committed, unconditionally, 25 to a // BatchWriteItem, retrying the items DynamoDB leaves unprocessed. func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time.Duration) error { diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 88bc3f76..65397b27 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -4,6 +4,7 @@ import ( "context" "errors" "fmt" + "maps" "net" "sync" "sync/atomic" @@ -155,6 +156,53 @@ func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { assert.Empty(t, claims[1].Token) } +// An SDK retry of a put whose first attempt was applied fails its condition +// on the put's own item: that is the caller's claim, not another request's. +func TestDynamo_RetriedPutKeepsItsOwnClaim(t *testing.T) { + t.Parallel() + var mu sync.Mutex + putTokens := map[string]string{} + var released []string + fake := &fakeDynamo{ + put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := idOf(in.Item[attrKey]) + mu.Lock() + putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) + mu.Unlock() + switch id { + case "k0": + return nil, &types.ConditionalCheckFailedException{Item: in.Item} + case "k1": + theirs := maps.Clone(in.Item) + theirs[attrToken] = &types.AttributeValueMemberB{Value: []byte("theirs")} + return nil, &types.ConditionalCheckFailedException{Item: theirs} + } + return nil, &types.InternalServerError{} + }, + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := idOf(in.Key[attrKey]) + mu.Lock() + defer mu.Unlock() + assert.Equal(t, putTokens[id], string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value)) + released = append(released, id) + return &dynamodb.DeleteItemOutput{}, nil + }, + } + _, m := openFakeWith(t, fake, DynamoConfig{Table: "dedupe", ReserveConcurrency: 1}) + + claims, err := m.Reserve(t.Context(), keys("k0", "k1"), time.Minute) + require.NoError(t, err) + assert.Equal(t, []Status{Claimed, InFlight}, []Status{claims[0].Status, claims[1].Status}, "a pending item is InFlight only under another token") + assert.Equal(t, putTokens["k0"], claims[0].Token) + + // k0 is sent and answered before k2 fails, so the undo owns it. + _, err = m.Reserve(t.Context(), keys("k0", "k2"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + mu.Lock() + defer mu.Unlock() + assert.Contains(t, released, "k0", "the failed Reserve's undo releases the retried put's own item") +} + func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { t.Parallel() var mu sync.Mutex diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go index 79bea18a..69be6b16 100644 --- a/tests/integration/dedupe_dynamodb_test.go +++ b/tests/integration/dedupe_dynamodb_test.go @@ -86,6 +86,23 @@ func awsError(status int, code string) *http.Response { const putItem = "DynamoDB_20120810.PutItem" +// landThenFail sends the first PutItem, then answers it 500 as if the +// response were lost: the write is applied and the SDK retries it. +type landThenFail struct { + next *http.Client + failed atomic.Bool +} + +func (l *landThenFail) Do(r *http.Request) (*http.Response, error) { + resp, err := l.next.Do(r) + if err != nil || r.Header.Get("X-Amz-Target") != putItem || !l.failed.CompareAndSwap(false, true) { + return resp, err + } + _, _ = io.Copy(io.Discard, resp.Body) + _ = resp.Body.Close() + return awsError(http.StatusInternalServerError, "InternalServerError"), nil +} + func TestDedupeDynamo_Conformance(t *testing.T) { t.Parallel() dedupetest.Run(t, func(t *testing.T) dedupetest.Harness { @@ -183,6 +200,35 @@ func TestDedupeDynamo_Throttled(t *testing.T) { } } +// The SDK's retry of an applied put fails its condition on the put's own +// item, which is still the caller's claim: without that, the id would be +// held InFlight for the lease by a claim nobody commits or releases. +func TestDedupeDynamo_RetriedPutKeepsItsClaim(t *testing.T) { + t.Parallel() + table := newDynamoTable() + lossy := &landThenFail{next: http.DefaultClient} + d := dynamoClient(t, table, dedupe.DynamoConfig{}, config.WithHTTPClient(lossy)) + require.NoError(t, d.CreateTable(t.Context())) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + peer := dynamoClient(t, table, dedupe.DynamoConfig{}).Tenant("acme") + require.NoError(t, peer.Apply(true)) + k := []dedupe.Key{{Table: "events", ID: "e1"}} + + claims, err := m.Reserve(t.Context(), k, time.Minute) + require.NoError(t, err) + require.True(t, lossy.failed.Load(), "the applied attempt was answered 500") + require.Equal(t, dedupe.Claimed, claims[0].Status) + other, err := peer.Reserve(t.Context(), k, time.Minute) + require.NoError(t, err) + assert.Equal(t, dedupe.InFlight, other[0].Status) + + require.NoError(t, m.Release(t.Context(), claims)) + other, err = peer.Reserve(t.Context(), k, time.Minute) + require.NoError(t, err) + assert.Equal(t, dedupe.Claimed, other[0].Status, "the claim's token was the applied put's, so Release freed the id") +} + func TestDedupeDynamo_Unreachable(t *testing.T) { t.Parallel() d, err := dedupe.NewDynamo(t.Context(), dedupe.DynamoConfig{ From 88958638bfa4cdcf4194a75dd6062845a8eb3b5a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:51:20 -0400 Subject: [PATCH 062/108] fix(dedupe): jittered DynamoDB retries that fit the call timeout With the retryer's MaxBackoff at 200 ms, the SDK's legacy backoff waited a fixed 200 ms before every retry, with no jitter: under the 250 ms default per-call timeout the third attempt was never sent, a throttled call ended as DeadlineExceeded with the throttle cause dropped, every put throttled together retried in lockstep, and a throttled Reserve took about twice the timeout once its undo hit the same schedule. The retryer now uses an explicit full-jitter backoff, uniform over [0, min(25 ms * 2^attempt, ceiling)], with the ceiling derived from the config as Timeout / (2 * (MaxAttempts - 1)). A call's retries therefore wait at most half its Timeout in all, whatever Timeout and MaxAttempts are set to, and a throttled call ends on its last attempt's answer: a max-attempts error carrying the throttle, mapped to ErrUnavailable. Commit's wait between rounds of unprocessed items is jittered the same way (25 ms doubling to 200 ms). Tests: the backoff's worst case fits half the Timeout across several configs and is jittered; a put throttled through the real SDK stack at the defaults ends as a MaxAttemptsError wrapping the throttle, inside the Timeout, after exactly MaxAttempts sends. Both failed on the old schedule. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 37 +++++++- internal/dedupe/dynamodb_test.go | 121 ++++++++++++++++++++++++++ 4 files changed, 156 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 792cb65d..babb8f4e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 9fb7dd8f..aeaed95a 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 5e71db3f..bd9a7485 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -6,6 +6,7 @@ import ( "errors" "fmt" "log/slog" + mathrand "math/rand/v2" "strconv" "sync" "time" @@ -45,7 +46,12 @@ const ( // commitRounds bounds the BatchWriteItem rounds one chunk gets before // its still-unprocessed items fail the Commit. commitRounds = 8 - tokenBytes = 16 + // commitBase and commitCeiling bound the jittered wait between Commit + // rounds; retryBase is the one the SDK retryer's backoff doubles from. + commitBase = 25 * time.Millisecond + commitCeiling = 200 * time.Millisecond + retryBase = 25 * time.Millisecond + tokenBytes = 16 // opReserve is the operation the breaker watches: Release and Commit // answers say nothing about whether a new Reserve would get through. @@ -64,7 +70,8 @@ type DynamoConfig struct { // only; it is also what unlocks CreateTable. Endpoint string // Timeout bounds each DynamoDB call, its SDK retries included. - // 0 = 250ms. + // 0 = 250ms. The retries' jittered backoff is capped so that together + // it waits at most half of Timeout (retryBackoff). Timeout time.Duration // MaxAttempts is the SDK retryer's attempts per call. 0 = 3. MaxAttempts int @@ -150,7 +157,7 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { standard := func(o *retry.StandardOptions) { o.MaxAttempts = cfg.MaxAttempts - o.MaxBackoff = 200 * time.Millisecond + o.Backoff = retryBackoff(cfg) } switch cfg.RetryMode { case "standard": @@ -165,6 +172,28 @@ func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { return nil, fmt.Errorf("dedupe: dynamodb retry_mode %q: want standard or adaptive", cfg.RetryMode) } +// retryBackoff is the SDK retryer's wait before a retry: full jitter, so +// puts throttled together do not retry in lockstep, under a ceiling that +// doubles from retryBase up to Timeout/(2·(MaxAttempts-1)). A call's retries +// then wait at most half its Timeout in all, so a throttled call ends on its +// last attempt's answer (ErrUnavailable, the throttle as its cause) unless +// the attempts themselves take the other half. +func retryBackoff(cfg DynamoConfig) retry.BackoffDelayerFunc { + ceiling := cfg.Timeout / time.Duration(2*max(cfg.MaxAttempts-1, 1)) + return func(attempt int, _ error) (time.Duration, error) { + return fullJitter(retryBase, ceiling, attempt), nil + } +} + +// fullJitter is uniform over [0, min(base·2^attempt, ceiling)]. +func fullJitter(base, ceiling time.Duration, attempt int) time.Duration { + d := min(base< Date: Sat, 26 Sep 2026 04:52:51 -0400 Subject: [PATCH 063/108] fix(dedupe): size the DynamoDB client's idle pool to its fan-out Reserve, Commit and Release fan out to ReserveConcurrency (64) calls at once, but the SDK's default transport keeps only 10 idle connections per host, so most of a wide Reserve's puts dialed a new connection every time. Against dynamodb-local, a warm 64-key Reserve opened 46 to 54 new connections; with the pool sized it reuses all 64. NewDynamo now passes a BuildableClient whose transport keeps ReserveConcurrency idle connections per host (and at least as many in all). It goes ahead of the caller's extra options, so a test's fault-injecting HTTP client still replaces it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 19 ++++++++++++++++--- internal/dedupe/dynamodb_test.go | 17 +++++++++++++++++ 4 files changed, 35 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index babb8f4e..118c4521 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index aeaed95a..8ee1ec6a 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index bd9a7485..bd15f84d 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -7,12 +7,14 @@ import ( "fmt" "log/slog" mathrand "math/rand/v2" + "net/http" "strconv" "sync" "time" "github.com/aws/aws-sdk-go-v2/aws" "github.com/aws/aws-sdk-go-v2/aws/retry" + awshttp "github.com/aws/aws-sdk-go-v2/aws/transport/http" "github.com/aws/aws-sdk-go-v2/config" "github.com/aws/aws-sdk-go-v2/service/dynamodb" "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" @@ -79,7 +81,8 @@ type DynamoConfig struct { // the client after throttles. RetryMode string // ReserveConcurrency bounds the parallel calls one Reserve, Commit or - // Release makes. 0 = 64. + // Release makes, and sizes the client's idle connection pool to match. + // 0 = 64. ReserveConcurrency int } @@ -128,7 +131,7 @@ type Dynamo struct { // NewDynamo builds the backend over a client from the SDK's default config // chain. extra is appended to the chain's options (a test's static -// credentials, say). It dials nothing: Check does. +// credentials or HTTP client, say). It dials nothing: Check does. func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.LoadOptions) error) (*Dynamo, error) { if cfg.Table == "" { return nil, errors.New("dedupe: dynamodb table is required") @@ -138,7 +141,7 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load if err != nil { return nil, err } - opts := []func(*config.LoadOptions) error{config.WithRetryer(retryer)} + opts := []func(*config.LoadOptions) error{config.WithRetryer(retryer), config.WithHTTPClient(newHTTPClient(cfg))} if cfg.Region != "" { opts = append(opts, config.WithRegion(cfg.Region)) } @@ -154,6 +157,16 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load return newDynamo(client, cfg), nil } +// newHTTPClient keeps an idle connection for every call one Reserve can have +// in flight: with the SDK's default of 10 per host, a wide Reserve would dial +// most of its puts afresh. +func newHTTPClient(cfg DynamoConfig) *awshttp.BuildableClient { + return awshttp.NewBuildableClient().WithTransportOptions(func(tr *http.Transport) { + tr.MaxIdleConnsPerHost = cfg.ReserveConcurrency + tr.MaxIdleConns = max(tr.MaxIdleConns, cfg.ReserveConcurrency) + }) +} + func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { standard := func(o *retry.StandardOptions) { o.MaxAttempts = cfg.MaxAttempts diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index d09af7b9..ddffeb48 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -16,6 +16,7 @@ import ( "github.com/aws/aws-sdk-go-v2/aws" "github.com/aws/aws-sdk-go-v2/aws/retry" + awshttp "github.com/aws/aws-sdk-go-v2/aws/transport/http" "github.com/aws/aws-sdk-go-v2/config" "github.com/aws/aws-sdk-go-v2/credentials" "github.com/aws/aws-sdk-go-v2/service/dynamodb" @@ -535,6 +536,22 @@ func TestDynamo_ThrottledCallEndsOnItsLastAttempt(t *testing.T) { assert.Equal(t, int64(calls*d.cfg.MaxAttempts), h.puts.Load()) } +// The idle pool holds a connection for every call one Reserve can have in +// flight, so a wide Reserve reuses them rather than dial; an HTTP client in +// extra (TestDynamo_ThrottledCallEndsOnItsLastAttempt's) replaces it. +func TestNewDynamo_SizesTheIdlePool(t *testing.T) { + t.Parallel() + for _, n := range []int{0, 8, 200} { + d, err := NewDynamo(t.Context(), DynamoConfig{Table: "dedupe", Region: "us-east-1", ReserveConcurrency: n}) + require.NoError(t, err) + client, ok := d.api.(*dynamodb.Client).Options().HTTPClient.(*awshttp.BuildableClient) + require.True(t, ok) + tr := client.GetTransport() + assert.GreaterOrEqual(t, tr.MaxIdleConnsPerHost, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) + assert.GreaterOrEqual(t, tr.MaxIdleConns, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) + } +} + func TestExpiresAt(t *testing.T) { t.Parallel() base := time.Unix(100, 0) From da1079ef31078988c27c128c54f55e3c097f4eab Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:53:51 -0400 Subject: [PATCH 064/108] test(dedupe): race Reserve against a Commit in progress Add a conformance case that spins Reserve on a Claimed key from both the factory's and the peer's client while its Commit runs, asserting no answer is ever Claimed and every answer is Duplicate once Commit returns. Also widen the existing concurrent-Reserve race from one key to 20 fresh keys, since a lock window narrow enough to survive one key can still slip past detection there. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/dedupetest/dedupetest.go | 91 ++++++++++++++++++++---- 1 file changed, 77 insertions(+), 14 deletions(-) diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index 5b9221b6..5eca743e 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -11,6 +11,7 @@ import ( "fmt" "strings" "sync" + "sync/atomic" "testing" "time" @@ -181,10 +182,15 @@ var cases = []struct { "retention 0 never expires") }}, {"concurrent reserves of one key claim it once", func(t *testing.T, s *suite) { - // #390: two requests carrying one id must not both publish. - const n = 64 + // #390: two requests carrying one id must not both publish. Run + // over many fresh keys, not just one: a race confined to a narrow + // lock window (e.g. one unlocked too early around a single + // backend read) can slip past a single key far more often than it + // triggers, so one key is a weak witness. + const n = 16 + const keys = 20 d, p := s.store(t, "acme"), s.peer(t, "acme") - race := func() []dedupe.Claim { + race := func(k dedupe.Key) []dedupe.Claim { out := make([]dedupe.Claim, n) var wg sync.WaitGroup for i := range n { @@ -193,7 +199,7 @@ var cases = []struct { store = p } wg.Go(func() { - c, err := store.Reserve(context.Background(), []dedupe.Key{key("e1")}, long) + c, err := store.Reserve(context.Background(), []dedupe.Key{k}, long) if assert.NoError(t, err) { out[i] = c[0] } @@ -202,18 +208,75 @@ var cases = []struct { wg.Wait() return out } - var winner []dedupe.Claim - for _, c := range race() { - if c.Status == dedupe.Claimed { - winner = append(winner, c) - } else { - assert.Equal(t, dedupe.InFlight, c.Status) + for round := range keys { + k := key(fmt.Sprintf("race-%d", round)) + var winner []dedupe.Claim + for _, c := range race(k) { + if c.Status == dedupe.Claimed { + winner = append(winner, c) + } else { + assert.Equal(t, dedupe.InFlight, c.Status, "round %d", round) + } + } + require.Len(t, winner, 1, "round %d: exactly one reserve claims the key", round) + require.NoError(t, d.Commit(t.Context(), winner, 0)) + for _, c := range race(k) { + assert.Equal(t, dedupe.Duplicate, c.Status, "round %d", round) } } - require.Len(t, winner, 1, "exactly one reserve claims the key") - require.NoError(t, d.Commit(t.Context(), winner, 0)) - for _, c := range race() { - assert.Equal(t, dedupe.Duplicate, c.Status) + }}, + {"a reserve racing a commit never claims, and settles to duplicate once it lands", func(t *testing.T, s *suite) { + // A backend's Commit must write the durable record before it drops + // the pending claim (e.g. the Pebble backend's synced batch flush + // ahead of releasing the in-memory claim) — a Reserve spinning + // against the same key during that window must never see Claimed, + // and must see Duplicate the instant Commit returns. Run over many + // fresh keys and both clients so a narrow unlock window isn't + // masked by luck on one key or one process's view. + const workers = 8 + const rounds = 20 + d, p := s.store(t, "acme"), s.peer(t, "acme") + for round := range rounds { + k := key(fmt.Sprintf("commit-race-%d", round)) + c := reserve(t, d, long, k) + var stop atomic.Bool + var claimed atomic.Int64 + var wg sync.WaitGroup + for i := range workers { + store := d + if i%2 == 1 { + store = p + } + wg.Go(func() { + for !stop.Load() { + got, err := store.Reserve(context.Background(), []dedupe.Key{k}, long) + if !assert.NoError(t, err) { + return + } + switch got[0].Status { + case dedupe.Claimed: + claimed.Add(1) + case dedupe.InFlight, dedupe.Duplicate: + default: + t.Errorf("round %d: unexpected status %v", round, got[0].Status) + } + } + }) + } + require.NoError(t, d.Commit(t.Context(), c, 0)) + stop.Store(true) + wg.Wait() + assert.Zero(t, claimed.Load(), "round %d: a reserve claimed a key mid-commit", round) + for i := range workers { + store := d + if i%2 == 1 { + store = p + } + got, err := store.Reserve(context.Background(), []dedupe.Key{k}, long) + if assert.NoError(t, err) { + assert.Equal(t, dedupe.Duplicate, got[0].Status, "round %d: reserve %d after commit", round, i) + } + } } }}, {"a key repeated in one call is claimed once", func(t *testing.T, s *suite) { From e52b336c3c7b41d3d66a608102093bda30f4c330 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:53:59 -0400 Subject: [PATCH 065/108] fix(dedupe): sweep holds the commit lock only to re-check and delete A sweep chunk held commitMu while it iterated to its 1,024th live key. Pebble skips point tombstones inside Next, so a chunk that started over a run of them (a tenant's version-0 block deleted by the first pass, or a table whose ids had all expired) held every Commit, and with it every deduped ingest response, for the whole run, on every pass until the run was compacted. The chunk now reads without the lock and collects the keys it would delete, then takes commitMu only to re-read each one and delete those still expired or version-0, without fsync as before. A Commit waits for at most 1,024 point reads and one unsynced batch, however many tombstones lie between the keys. The mid-chunk test splits in two, one per gap a Commit can land in: after the unlocked read, where the re-read keeps the key, and after the re-read, where the lock holds the Commit off until the delete is done. Each fails when its guard is removed. A new test races Commits against a chunk that starts over 300,000 tombstones: under the race detector the old chunk kept one waiting 241-278 ms, the new one at most 11 ms. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 2 +- internal/dedupe/embedded.go | 13 +-- internal/dedupe/sweep.go | 119 ++++++++++++++++++-------- internal/dedupe/sweep_test.go | 87 ++++++++++++++++--- 6 files changed, 172 insertions(+), 53 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7b7935f0..9c9feadb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,7 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. -- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains an optional **`dedupe.retention`** key, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. A `config.json` without the key keeps ids forever, and a table override without one inherits the tenant's, so an existing directory needs no change. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains an optional **`dedupe.retention`** key, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. A `config.json` without the key keeps ids forever, and a table override without one inherits the tenant's, so an existing directory needs no change. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk without a lock, then re-reads the expired and version-0 ones under a lock `Commit` also takes and deletes, without fsync, those that still are, so an id committed again after the sweep read it is never deleted, and a `Commit` waits for at most one chunk's re-reads, never for the deleted keys a chunk steps over. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 642d25a0..b7fa3db5 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -125,7 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose records were definitely not published (a refused or never-sent publish; one whose outcome is unknown is left to lapse instead). A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. -- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync, each value carrying its expiry (`0` = never), which `Reserve` honors on read. A background sweep (`sweep.go`), started when the instance opens and stopped before it closes, deletes expired keys and the version-0 keys from before the table joined the key — told apart by their value, which is never a current commit's, since a bare id from before tenants led the key could spell a current one ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)): a minute after opening, then hourly, 1,024 keys per chunk, holding a lock `Commit` also takes, so a key re-committed after the sweep read it is never deleted; `wavehouse_dedupe_swept_keys_total{reason}` counts what it deletes. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync, each value carrying its expiry (`0` = never), which `Reserve` honors on read. A background sweep (`sweep.go`), started when the instance opens and stopped before it closes, deletes expired keys and the version-0 keys from before the table joined the key — told apart by their value, which is never a current commit's, since a bare id from before tenants led the key could spell a current one ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)): a minute after opening, then hourly, 1,024 keys per chunk, read without a lock, so the deleted keys a chunk steps over (Pebble keeps them until it compacts) never hold up a `Commit`, then re-read under a lock `Commit` also takes and deleted only if still expired or version-0, so a key re-committed after the sweep read it is never deleted; `wavehouse_dedupe_swept_keys_total{reason}` counts what it deletes. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 32e9f5b2..48fc8817 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads 1,024 keys at a time, deleting the expired ones, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. +With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path. It reads 1,024 keys at a time without holding up commits, then re-reads the expired ones and deletes those still expired; a commit waits only for that last step, at most 1,024 point reads and one unsynced write, however many deleted keys the read stepped over. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index d0fe87ac..5eece3f3 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -35,8 +35,9 @@ type Embedded struct { open int // tenant stores open over db stopSweep func() // stops db's sweep - // commitMu is read-held by Commit and held by a sweep chunk, so a sweep - // never deletes a key a Commit rewrote after the sweep read it. + // commitMu is read-held by Commit and held by a sweep chunk while it + // re-reads and deletes, so a sweep never deletes a key a Commit rewrote + // after the sweep read it. commitMu sync.RWMutex sweepFirst time.Duration sweepEvery time.Duration @@ -47,9 +48,11 @@ type Embedded struct { // readHook, when set, runs before each Pebble read in Reserve; a test // makes it fail to exercise Reserve's all-or-nothing error path. readHook func() error - // sweepHook, when set, runs in a sweep chunk between reading its keys and - // deleting them; a test races a Commit into that gap. - sweepHook func() + // sweepScanHook and sweepDeleteHook, when set, run in a sweep chunk: + // between its unlocked read and its re-read, and between its re-read and + // its delete. A test races a Commit into each gap. + sweepScanHook func() + sweepDeleteHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index 70018684..f1c4ca18 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -3,6 +3,7 @@ package dedupe import ( "bytes" "context" + "errors" "fmt" "log/slog" "time" @@ -20,9 +21,9 @@ import ( const ( sweepInterval = time.Hour sweepFirstDelay = time.Minute - // sweepChunk keys are read and deleted per lock hold, with sweepPause - // between chunks: at most ~100k keys a second, and a Commit never waits - // longer than one chunk. + // sweepChunk keys are read per chunk, with sweepPause between chunks: at + // most ~100k keys a second. A Commit waits only for a chunk's re-reads + // and deletes, never for its read, however many tombstones it skips. sweepChunk = 1024 sweepPause = 10 * time.Millisecond ) @@ -98,21 +99,41 @@ func (e *Embedded) sweep(ctx context.Context, db *pebble.DB) (sweepResult, error } // sweepChunk deletes the sweepable keys among the next sweepChunk keys from -// from, returning where the next chunk starts (nil at the end). It holds -// commitMu, so no Commit lands between reading a key and deleting it: a key -// re-committed after it expired is never deleted with its new value. +// from, returning where the next chunk starts (nil at the end). It reads them +// without commitMu, since Pebble skips the tombstones between keys inside the +// read and a run of them left by an earlier pass would otherwise hold every +// Commit for its whole length. func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, res *sweepResult) ([]byte, error) { - e.commitMu.Lock() - defer e.commitMu.Unlock() - now := e.now() + candidates, next, err := sweepCandidates(db, from, e.now()) + if err != nil || len(candidates) == 0 { + return next, err + } + if e.sweepScanHook != nil { + e.sweepScanHook() + } + expired, v0, err := e.deleteSweepable(db, candidates) + if err != nil { + return nil, err + } + res.Expired += int(expired) + res.Version0 += int(v0) + if expired > 0 { + sweptKeysCounter.Add(ctx, expired, metric.WithAttributes(attribute.String(sweptAttribute, sweptExpired))) + } + if v0 > 0 { + sweptKeysCounter.Add(ctx, v0, metric.WithAttributes(attribute.String(sweptAttribute, sweptVersion0))) + } + return next, nil +} + +// sweepCandidates reads the next sweepChunk keys, starting at from, and +// returns those sweepable at now and where the next chunk starts (nil at the +// end). +func sweepCandidates(db *pebble.DB, from []byte, now time.Time) (candidates [][]byte, next []byte, err error) { it, err := db.NewIter(&pebble.IterOptions{LowerBound: from}) if err != nil { - return nil, fmt.Errorf("dedupe sweep: %w", err) + return nil, nil, fmt.Errorf("dedupe sweep: %w", err) } - b := db.NewBatch() - defer func() { _ = b.Close() }() - var next []byte - var expired, v0 int64 seen := 0 for valid := it.First(); valid; valid = it.Next() { if seen == sweepChunk { @@ -120,38 +141,66 @@ func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, r break } seen++ - k := it.Key() - val := it.Value() - switch { - case !isCommit(val): + if sweepReason(it.Value(), now) != "" { + candidates = append(candidates, bytes.Clone(it.Key())) + } + } + if err := it.Close(); err != nil { + return nil, nil, fmt.Errorf("dedupe sweep: %w", err) + } + return candidates, next, nil +} + +// deleteSweepable re-reads each candidate and deletes those still sweepable, +// holding commitMu so no Commit lands between the re-read and the delete: a +// key re-committed after the unlocked read is never deleted with its new +// value. +func (e *Embedded) deleteSweepable(db *pebble.DB, candidates [][]byte) (expired, v0 int64, err error) { + e.commitMu.Lock() + defer e.commitMu.Unlock() + now := e.now() + b := db.NewBatch() + defer func() { _ = b.Close() }() + for _, k := range candidates { + val, closer, err := db.Get(k) + if errors.Is(err, pebble.ErrNotFound) { + continue + } + if err != nil { + return 0, 0, fmt.Errorf("dedupe sweep: %w", err) + } + reason := sweepReason(val, now) + _ = closer.Close() + switch reason { + case sweptVersion0: v0++ - case committedExpired(val, now): + case sweptExpired: expired++ default: continue } if err := b.Delete(k, nil); err != nil { - _ = it.Close() - return nil, fmt.Errorf("dedupe sweep: %w", err) + return 0, 0, fmt.Errorf("dedupe sweep: %w", err) } } - if err := it.Close(); err != nil { - return nil, fmt.Errorf("dedupe sweep: %w", err) - } - if e.sweepHook != nil { - e.sweepHook() + if e.sweepDeleteHook != nil { + e.sweepDeleteHook() } // NoSync: a delete lost to a crash is redone by the next pass. if err := b.Commit(pebble.NoSync); err != nil { - return nil, fmt.Errorf("dedupe sweep: %w", err) - } - res.Expired += int(expired) - res.Version0 += int(v0) - if expired > 0 { - sweptKeysCounter.Add(ctx, expired, metric.WithAttributes(attribute.String(sweptAttribute, sweptExpired))) + return 0, 0, fmt.Errorf("dedupe sweep: %w", err) } - if v0 > 0 { - sweptKeysCounter.Add(ctx, v0, metric.WithAttributes(attribute.String(sweptAttribute, sweptVersion0))) + return expired, v0, nil +} + +// sweepReason is why the sweep deletes a key holding val at now, or "" when +// it keeps it. +func sweepReason(val []byte, now time.Time) string { + switch { + case !isCommit(val): + return sweptVersion0 + case committedExpired(val, now): + return sweptExpired } - return next, nil + return "" } diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index da4ec951..0c7ff46e 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -100,28 +100,52 @@ func TestEmbedded_SweepDeletesExpiredAndVersionZeroKeys(t *testing.T) { assert.Equal(t, sweepResult{}, res, "a second pass finds nothing") } -// A Commit that arrives while a sweep chunk has read an expired key but not -// yet deleted it waits for the chunk, so the new commit is never deleted with -// the old value. Without the lock the Commit lands in the gap and the sweep -// then deletes it; the wait below only ever lets that pass, never fail. -func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { - t.Parallel() +// expiredAndClaimed commits id "e1" with an hour's retention, lets it expire +// and claims it again, for a test to commit mid-sweep. +func expiredAndClaimed(t *testing.T) (*Embedded, *Managed, []Claim) { + t.Helper() e := NewEmbedded(t.TempDir()) clock := newStepClock() SetClock(e, clock.now) m := switchedOn(t, e, "acme") commitIDs(t, m, time.Hour, "e1") clock.advance(2 * time.Hour) - claims, err := m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) require.NoError(t, err) require.Equal(t, Claimed, claims[0].Status, "expired: claimable again") + return e, m, claims +} + +// A key committed again after a sweep chunk read it as expired, but before +// the chunk re-read it, is kept: the re-read sees the new commit. +func TestEmbedded_SweepKeepsAKeyCommittedAfterItsRead(t *testing.T) { + t.Parallel() + e, m, claims := expiredAndClaimed(t) + var commitErr error + e.sweepScanHook = func() { commitErr = m.Commit(context.Background(), claims, time.Hour) } + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + require.NoError(t, commitErr) + assert.Equal(t, sweepResult{}, res) + + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.True(t, dup, "the commit made after the read survived the sweep") +} + +// A Commit that arrives while a sweep chunk has re-read an expired key but +// not yet deleted it waits for the chunk, so the new commit is never deleted +// with the old value. Without the lock the Commit lands in the gap and the +// sweep then deletes it; the wait below only ever lets that pass, never fail. +func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { + t.Parallel() + e, m, claims := expiredAndClaimed(t) done := make(chan error, 1) - e.sweepHook = func() { + e.sweepDeleteHook = func() { go func() { done <- m.Commit(context.Background(), claims, time.Hour) }() select { - case <-done: - done <- nil + case err := <-done: + done <- err case <-time.After(50 * time.Millisecond): } } @@ -135,6 +159,49 @@ func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { assert.True(t, dup, "the commit made mid-chunk survived the sweep") } +// A chunk that starts over a long run of tombstones, as a tenant's version-0 +// block leaves until Pebble compacts it, reads through the run without the +// lock, so a Commit racing it waits for the chunk's re-reads and deletes +// alone. Sized for the race detector, which the unit suite runs under: there, +// a chunk holding the lock across this run kept a Commit waiting ~250 ms. +func TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + m := switchedOn(t, e, "acme") + b := e.db.NewBatch() + for i := range 300_000 { + require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%06d", i), nil)) + } + // A version-0 key after the run, so the chunk has one to delete. + require.NoError(t, b.Set([]byte("acme\x01"), make([]byte, 8), nil)) + require.NoError(t, b.Commit(pebble.NoSync)) + require.NoError(t, e.db.Flush()) + + type swept struct { + res sweepResult + err error + } + done := make(chan swept, 1) + go func() { + res, err := e.sweep(context.Background(), e.db) + done <- swept{res, err} + }() + var slowest time.Duration + for i := 0; ; i++ { + select { + case s := <-done: + require.NoError(t, s.err) + require.Equal(t, sweepResult{Version0: 1}, s.res) + assert.Less(t, slowest, 100*time.Millisecond, "the slowest Commit racing the sweep") + return + default: + } + start := time.Now() + commitIDs(t, m, 0, fmt.Sprintf("c%d", i)) + slowest = max(slowest, time.Since(start)) + } +} + // A retention is honoured on read before any sweep has run: the key is a // duplicate until the retention ends and claimable from that instant. func TestEmbedded_RetentionHonouredOnRead(t *testing.T) { From ffa4dd6875c168154e68fd4219601d7b5a380ea3 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:54:27 -0400 Subject: [PATCH 066/108] docs(settings): name dedupe.retention in the last no-defaults claims dedupe.retention is the one config.json key the binary defaults (missing means "0", forever). The ingest handler's dedupe comment, the seed's doc comment, a registry test comment and the settings-directory CHANGELOG entry still said there were none. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- internal/api/ingest.go | 4 ++-- internal/settings/registry_test.go | 2 +- internal/settings/seed.go | 3 ++- 4 files changed, 6 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9c9feadb..8019c2e3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -24,7 +24,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Schema discovery captures each table's DDL, its columns' ordinals and default expressions, and the server version** (`internal/discovery/discovery.go`, `internal/testutil/testutil.go`): `Column` gains `DefaultExpression` and `Position` (both from a widened `system.columns` select), `TableSchema` gains `DDL` from `system.tables.create_table_query`, and `SchemaRegistry` gains `ServerVersion()` from a `SELECT version()` probe next to the existing `SELECT timezone()`. Groundwork for the native type layer, captured on the same refresh as the columns so a stale version cannot outlive the schemas it describes. That is a publication guarantee, not a same-server one: `chconn.Manager` resolves the connection per call, so a reload changing `clickhouse.addr` mid-refresh can still pair a version from one server with schemas from another — narrow, and self-correcting on the next refresh. `DDL` is `json:"-"` and does **not** appear in `/v1/ops/schema`: that endpoint marshals `TableSchema` straight to the client, and an external-engine table (S3, MySQL, PostgreSQL, Kafka) renders its wiring there unconditionally — endpoint, bucket or host, database, username, S3 access key id. ClickHouse masks the password itself as `[HIDDEN]` from ~23.9 (verified on 26.7.3), so the exposure is the topology rather than the secret — except on an older server, or one with `display_secrets_in_show_and_select` enabled. `position` and `default_expression` are additive fields in the response. A table listed in `system.tables` with no `system.columns` rows is skipped rather than published column-less, and both new queries fail the refresh on error exactly as `timezone()` and `system.columns` do — callers keep the prior cache and retry. -- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the sweeper purges it back under the limit, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. +- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the sweeper purges it back under the limit, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** but one — every `config.json` key except `dedupe.retention` (missing means `"0"`, forever) is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. - **"Was this page helpful?" feedback widget on every docs page** (`docs/src/components/PageFeedback.astro` (new), `docs/src/components/Footer.astro`): a thumbs-up / thumbs-down vote below the page content, captured to PostHog as `docs_feedback` with `{ helpful, page }`. It renders from `Footer.astro`'s sidebar branch — the same indirection the Cloud CTA uses — rather than a per-page import or frontmatter flag, so every content page gets it automatically, including ones not written yet; it sits *below* the Cloud CTA on the pages that carry one, and splash pages (the homepage and 404) take the other footer branch and never render it. One vote per page per visitor: the choice is remembered in `localStorage` keyed by pathname, and a revisit renders the thanks message instead of re-prompting (storage is a nicety, not the record — a browser with storage disabled still votes). - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 1189547a..58126f94 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -704,8 +704,8 @@ func (h *IngestHandler) prepareRecord( // Optional deduplication. The dedupe settings resolve per record // from one snapshot (table override → global; the settings directory - // always states them, so no compiled fallback is needed), so a reload - // lands at a record boundary. A Deduplicator without a settings source is + // states them all but dedupe.retention, whose absence means "0"), so a + // reload lands at a record boundary. A Deduplicator without a settings source is // a wiring bug, not a mode — main wires both or neither. The id is claimed // in ingestWindow, once every record of the window is encoded, so nothing // but the publish can fail while the claim is held. diff --git a/internal/settings/registry_test.go b/internal/settings/registry_test.go index 3ac222cc..a5abce84 100644 --- a/internal/settings/registry_test.go +++ b/internal/settings/registry_test.go @@ -79,7 +79,7 @@ func TestRegistry_ReloadWithWarningsAdopts(t *testing.T) { // TestOpen_RejectsInvalid pins the boot contract: an invalid directory yields // no Registry at all — there is no "store without a document" state and no -// compiled defaults to fall back on. +// compiled default for a required key to fall back on. func TestOpen_RejectsInvalid(t *testing.T) { t.Parallel() files := validFiles() diff --git a/internal/settings/seed.go b/internal/settings/seed.go index d9e979ae..5315c173 100644 --- a/internal/settings/seed.go +++ b/internal/settings/seed.go @@ -10,7 +10,8 @@ import ( // seedFS holds the starter settings directory: every file present, every // key set to its default. The checked-in seed/ directory is the ONE place -// defaults live — the binary has no compiled fallbacks. Its one consumer is +// defaults live — the binary's one compiled fallback is dedupe.retention +// (missing means "0", forever). Its one consumer is // this embed, so `wavehouse bootstrap` can write the directory anywhere // without a source tree; the container images ship no settings (the operator // mounts or seeds /app/settings), same as they ship no policy file. From eb5dea9035b8daf76ecbdb13f0e9cef70e1befb3 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:55:03 -0400 Subject: [PATCH 067/108] fix(dedupe): retry a throttled DynamoDB commit like unprocessed items commitChunk returned on the first ErrUnavailable from BatchWriteItem. DynamoDB raises the throttle exception only when it processed none of the batch, so a total throttle got the SDK's retries inside one call while a partial throttle got eight backoff rounds: Commit gave up sooner under heavier throttling. These records are already published, and a lost commit lets a retry publish them again. A transient (ErrUnavailable-class) error from the call now counts as a round in which every item came back unprocessed: the same jittered backoff, the same eight-round cap, and the last round's error (with its cause) if the cap is reached. The puts are idempotent, so resending a batch whose fate a timeout left unknown is safe. Non-transient errors stay final. The unprocessed-items metric still counts only what DynamoDB reported unprocessed. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 23 +++++++++++--- internal/dedupe/dynamodb_test.go | 44 +++++++++++++++++++++++++++ 4 files changed, 64 insertions(+), 7 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 118c4521..671b8bd3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 8ee1ec6a..b8ab38cc 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index bd15f84d..f5c22b59 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -46,7 +46,7 @@ const ( // batchWriteMax is BatchWriteItem's per-call item limit. batchWriteMax = 25 // commitRounds bounds the BatchWriteItem rounds one chunk gets before - // its still-unprocessed items fail the Commit. + // its still-unprocessed (or still-throttled) items fail the Commit. commitRounds = 8 // commitBase and commitCeiling bound the jittered wait between Commit // rounds; retryBase is the one the SDK retryer's backoff doubles from. @@ -125,7 +125,7 @@ type Dynamo struct { breaker *breaker metrics dynamoMetrics // commitBackoff is the wait before retrying the attempt'th round of - // unprocessed items. + // unprocessed or throttled items. commitBackoff func(attempt int) time.Duration } @@ -419,7 +419,8 @@ func heldBy(item map[string]types.AttributeValue, token string) bool { } // Commit overwrites every claim's item as committed, unconditionally, 25 to a -// BatchWriteItem, retrying the items DynamoDB leaves unprocessed. +// BatchWriteItem, retrying the items DynamoDB leaves unprocessed and a batch +// that failed transiently. func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time.Duration) error { if len(claims) == 0 { return nil @@ -468,16 +469,28 @@ func (s *dynamoStore) commitChunk(ctx context.Context, writes []types.WriteReque } return err }) - if err != nil { + switch { + case err == nil: + case errors.Is(err, ErrUnavailable): + // DynamoDB throttles a batch whole only when it processed none of + // it, and a timeout leaves its fate unknown: retry it whole, as a + // round that left every item unprocessed. The puts are idempotent. + unprocessed = writes + default: return err } if len(unprocessed) == 0 { return nil } if attempt+1 >= commitRounds { + if err != nil { + return err + } return fmt.Errorf("%w: dynamodb batch_write_item: %d items still unprocessed", ErrUnavailable, len(unprocessed)) } - s.d.metrics.unprocessed.Add(ctx, int64(len(unprocessed))) + if err == nil { + s.d.metrics.unprocessed.Add(ctx, int64(len(unprocessed))) + } writes = unprocessed select { case <-ctx.Done(): diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index ddffeb48..31f54f85 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -303,6 +303,50 @@ func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { require.ErrorIs(t, err, ErrUnavailable) } +// DynamoDB answers a BatchWriteItem it processed none of with a throttle, +// not with every item unprocessed: the chunk gets the same rounds either way. +func TestDynamo_CommitRetriesAFailedBatch(t *testing.T) { + t.Parallel() + throttle := &types.ProvisionedThroughputExceededException{} + for _, tc := range []struct { + name string + fails int + err error + calls int64 + committed bool + unavailable bool + }{ + {"throttled, then through", 3, throttle, 4, true, false}, + {"throttled every round", 1 << 10, throttle, commitRounds, false, true}, + {"timed out, then through", 1, fmt.Errorf("op: %w", context.DeadlineExceeded), 2, true, false}, + {"a configuration bug is final", 1 << 10, &types.ResourceNotFoundException{}, 1, false, false}, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + var calls atomic.Int64 + _, m := openFake(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + assert.Len(t, in.RequestItems["dedupe"], 2, "the whole batch, every round") + if calls.Add(1) <= int64(tc.fails) { + return nil, tc.err + } + return &dynamodb.BatchWriteItemOutput{}, nil + }}) + var claims []Claim + for _, k := range keys("a", "b") { + claims = append(claims, Claim{Key: k, Status: Claimed, Token: "t"}) + } + err := m.Commit(t.Context(), claims, 0) + assert.Equal(t, tc.calls, calls.Load()) + if tc.committed { + require.NoError(t, err) + return + } + require.ErrorIs(t, err, tc.err, "the last round's cause is kept") + assert.Equal(t, tc.unavailable, errors.Is(err, ErrUnavailable)) + }) + } +} + func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { t.Parallel() _, m := openFake(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { From f40c76aa4acba1ebea48439f86a5584732353a17 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:55:18 -0400 Subject: [PATCH 068/108] fix(config): dedupe defaults in defaults(); refuse an explicit zero The dedupe keys carried cleanenv env-default tags, which main no longer allows: cleanenv applies them after the YAML decode to any field still zero, so an explicit zero in config.yaml silently became the default. dedupe.lease, reserve_concurrency and the dynamodb block's timeout, max_attempts and retry_mode now default in defaults(), and a zero lease, concurrency, timeout or attempt count, or an empty retry_mode, refuses boot (the dynamodb block's only while dynamodb is selected). Each is a refusedZeros case; the docs-defaults test gains a duration parser. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 12 ++--- internal/config/backends.go | 30 ++++++------- internal/config/backends_test.go | 17 +++---- internal/config/config.go | 8 +++- internal/config/defaults_test.go | 59 +++++++++++++++++-------- 6 files changed, 77 insertions(+), 51 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2d5b586b..db2333b6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index e98b69ae..7dbd3171 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -56,8 +56,8 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` = the default. | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` = the default. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` refuses boot. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe @@ -65,12 +65,12 @@ Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(none)* | The shared table. Required. | +| `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(required)* | The shared table. | | `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`); no region from either refuses boot. | | `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | -| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. `0` = the default. | -| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. `0` = the default. | -| `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. `0` refuses boot. | +| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. `0` refuses boot. | +| `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. Anything else, empty included, refuses boot. | | `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | ### Process roles diff --git a/internal/config/backends.go b/internal/config/backends.go index 6286956b..493e9488 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -75,10 +75,10 @@ type Dedupe struct { Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND"` // Lease is how long a claimed id stays pending while its record is // published; a claim its request never settles lapses after it. - Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE" env-default:"30s"` + Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE"` // ReserveConcurrency bounds the parallel calls one Reserve, Commit or // Release makes to a remote backend. Pebble ignores it. - ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY" env-default:"64"` + ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY"` DynamoDB DedupeDynamoDBConfig `yaml:"dynamodb"` } @@ -93,23 +93,23 @@ type DedupeDynamoDBConfig struct { Region string `yaml:"region" env:"WH_DEDUPE_DYNAMODB_REGION"` // Endpoint points the client at dynamodb-local. Endpoint string `yaml:"endpoint" env:"WH_DEDUPE_DYNAMODB_ENDPOINT"` - Timeout time.Duration `yaml:"timeout" env:"WH_DEDUPE_DYNAMODB_TIMEOUT" env-default:"250ms"` - MaxAttempts int `yaml:"max_attempts" env:"WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS" env-default:"3"` - RetryMode string `yaml:"retry_mode" env:"WH_DEDUPE_DYNAMODB_RETRY_MODE" env-default:"standard"` + Timeout time.Duration `yaml:"timeout" env:"WH_DEDUPE_DYNAMODB_TIMEOUT"` + MaxAttempts int `yaml:"max_attempts" env:"WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS"` + RetryMode string `yaml:"retry_mode" env:"WH_DEDUPE_DYNAMODB_RETRY_MODE"` // CreateTable creates the table at boot if it is missing. Development // only: refused unless Endpoint is set. - CreateTable bool `yaml:"create_table" env:"WH_DEDUPE_DYNAMODB_CREATE_TABLE" env-default:"false"` + CreateTable bool `yaml:"create_table" env:"WH_DEDUPE_DYNAMODB_CREATE_TABLE"` } func (d Dedupe) validate() error { if err := checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends); err != nil { return err } - if d.Lease < 0 { - return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) must be >= 0, got %s", d.Lease) + if d.Lease <= 0 { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) must be > 0, got %s", d.Lease) } - if d.ReserveConcurrency < 0 { - return fmt.Errorf("dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be >= 0, got %d", d.ReserveConcurrency) + if d.ReserveConcurrency <= 0 { + return fmt.Errorf("dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0, got %d", d.ReserveConcurrency) } if d.Backend == DedupeDynamoDB { return d.DynamoDB.validate() @@ -121,11 +121,11 @@ func (d DedupeDynamoDBConfig) validate() error { switch { case strings.TrimSpace(d.Table) == "": return errors.New("dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required when dedupe.backend is dynamodb") - case d.Timeout < 0: - return fmt.Errorf("dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be >= 0, got %s", d.Timeout) - case d.MaxAttempts < 0: - return fmt.Errorf("dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be >= 0, got %d", d.MaxAttempts) - case d.RetryMode != "" && d.RetryMode != "standard" && d.RetryMode != "adaptive": + case d.Timeout <= 0: + return fmt.Errorf("dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be > 0, got %s", d.Timeout) + case d.MaxAttempts <= 0: + return fmt.Errorf("dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be > 0, got %d", d.MaxAttempts) + case d.RetryMode != "standard" && d.RetryMode != "adaptive": return fmt.Errorf("dedupe.dynamodb.retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) %q: want standard or adaptive", d.RetryMode) case d.CreateTable && d.Endpoint == "": return errors.New("dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only: set dedupe.dynamodb.endpoint, or create the table with your infrastructure code") diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 26144915..56f0d083 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -11,11 +11,11 @@ import ( ) // withDefaultBackends sets what defaults() would: a literal Config -// names no backend and no role, and Validate refuses that. +// names no backend, no role and no dedupe lease, and Validate refuses that. func withDefaultBackends(c Config) *Config { c.Roles = AllRoles() c.MQ.Backend, c.Cache.Backend = MQEmbedded, CacheLocal - c.Dedupe.Backend, c.Coord.Backend = DedupePebble, CoordLocal + c.Dedupe, c.Coord.Backend = defaults().Dedupe, CoordLocal return &c } @@ -254,15 +254,11 @@ func TestValidate_Dedupe(t *testing.T) { want string // "" = valid }{ {"dynamodb", dynamo, ""}, - {"zero values read as the defaults", func(c *Config) { - c.Dedupe.Backend = DedupeDynamoDB - c.Dedupe.DynamoDB = DedupeDynamoDBConfig{Table: "t"} - }, ""}, {"create_table with an endpoint", func(c *Config) { dynamo(c) c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true }, ""}, - {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB.CreateTable = true }, ""}, + {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB = DedupeDynamoDBConfig{CreateTable: true} }, ""}, {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, ""}, {"create_table without an endpoint", func(c *Config) { dynamo(c) @@ -270,9 +266,14 @@ func TestValidate_Dedupe(t *testing.T) { }, "dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only"}, {"no table", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Table = " " }, "dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required"}, {"retry mode", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.RetryMode = "legacy" }, `retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) "legacy"`}, + {"zero timeout", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Timeout = 0 }, "dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be > 0"}, {"negative timeout", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Timeout = -time.Second }, "dedupe.dynamodb.timeout"}, + {"zero attempts", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.MaxAttempts = 0 }, "dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be > 0"}, {"negative attempts", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.MaxAttempts = -1 }, "dedupe.dynamodb.max_attempts"}, - {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be >= 0"}, + {"no retry mode", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.RetryMode = "" }, `retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) ""`}, + {"zero lease", func(c *Config) { c.Dedupe.Lease = 0 }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0, got 0s"}, + {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0"}, + {"zero concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = 0 }, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0"}, {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, {"lease past the duplicate window", func(c *Config) { c.Dedupe.Lease = 3 * time.Minute }, "exceeds the embedded mq's 2m0s duplicate window"}, } diff --git a/internal/config/config.go b/internal/config/config.go index 95bfacb1..fcd0cd73 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -7,6 +7,7 @@ import ( "os" "slices" "strings" + "time" "github.com/ilyakaznacheev/cleanenv" ) @@ -256,8 +257,11 @@ func defaults() Config { Server: Server{Port: 8080, ShutdownTimeout: 10}, MQ: MQ{Backend: MQEmbedded}, Cache: Cache{Backend: CacheLocal, L1MaxCost: 64 << 20}, - Dedupe: Dedupe{Backend: DedupePebble}, - Coord: Coord{Backend: CoordLocal}, + Dedupe: Dedupe{ + Backend: DedupePebble, Lease: 30 * time.Second, ReserveConcurrency: 64, + DynamoDB: DedupeDynamoDBConfig{Timeout: 250 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard"}, + }, + Coord: Coord{Backend: CoordLocal}, OTel: OTel{ Traces: OTelTraces{Enabled: true, SampleRate: 1.0}, Metrics: OTelMetrics{Enabled: true}, diff --git a/internal/config/defaults_test.go b/internal/config/defaults_test.go index 6877be4b..cbbfdef8 100644 --- a/internal/config/defaults_test.go +++ b/internal/config/defaults_test.go @@ -9,6 +9,7 @@ import ( "strconv" "strings" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -41,35 +42,53 @@ var zeroCases = []zeroCase{ // refusedZeros are the non-zero defaults whose zero Validate refuses: written // in the file, the zero must reach Validate rather than become the default. +// also holds the keys a sub-block's zero needs to be read at all. var refusedZeros = []struct { key string zero any err string + also map[string]any }{ - {"server.port", 0, "server.port 0 out of range"}, - {"mq.backend", "", `mq.backend (WH_MQ_BACKEND) ""`}, - {"cache.backend", "", `cache.backend (WH_CACHE_BACKEND) ""`}, - {"dedupe.backend", "", `dedupe.backend (WH_DEDUPE_BACKEND) ""`}, - {"coord.backend", "", `coord.backend (WH_COORD_BACKEND) ""`}, - {"roles", []string{}, "roles (WH_ROLES) is empty"}, + {"server.port", 0, "server.port 0 out of range", nil}, + {"mq.backend", "", `mq.backend (WH_MQ_BACKEND) ""`, nil}, + {"cache.backend", "", `cache.backend (WH_CACHE_BACKEND) ""`, nil}, + {"dedupe.backend", "", `dedupe.backend (WH_DEDUPE_BACKEND) ""`, nil}, + {"coord.backend", "", `coord.backend (WH_COORD_BACKEND) ""`, nil}, + {"roles", []string{}, "roles (WH_ROLES) is empty", nil}, + {"dedupe.lease", "0s", "dedupe.lease (WH_DEDUPE_LEASE) must be > 0", nil}, + {"dedupe.reserve_concurrency", 0, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0", nil}, + {"dedupe.dynamodb.timeout", "0s", "dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be > 0", dynamoSelected}, + {"dedupe.dynamodb.max_attempts", 0, "dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be > 0", dynamoSelected}, + {"dedupe.dynamodb.retry_mode", "", `dedupe.dynamodb.retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) ""`, dynamoSelected}, } -// yamlAt renders a file setting key to value, plus otel.enabled: true so -// the test can tell the file was read. -func yamlAt(t *testing.T, key string, value any) string { +// dynamoSelected is what the dedupe.dynamodb block needs to be read. +var dynamoSelected = map[string]any{"dedupe.backend": "dynamodb", "dedupe.dynamodb.table": "t"} + +// yamlAt renders a file setting key to value, and each dotted key of also to +// its value, plus otel.enabled: true so the test can tell the file was read. +func yamlAt(t *testing.T, key string, value any, also ...map[string]any) string { t.Helper() tree := map[string]any{"otel": map[string]any{"enabled": true}} - node := tree - parts := strings.Split(key, ".") - for _, p := range parts[:len(parts)-1] { - sub, ok := node[p].(map[string]any) - if !ok { - sub = map[string]any{} - node[p] = sub + set := func(key string, value any) { + node := tree + parts := strings.Split(key, ".") + for _, p := range parts[:len(parts)-1] { + sub, ok := node[p].(map[string]any) + if !ok { + sub = map[string]any{} + node[p] = sub + } + node = sub + } + node[parts[len(parts)-1]] = value + } + for _, m := range also { + for k, v := range m { + set(k, v) } - node = sub } - node[parts[len(parts)-1]] = value + set(key, value) out, err := yaml.Marshal(tree) require.NoError(t, err) return string(out) @@ -136,7 +155,7 @@ func TestLoad_YAMLZeroIsRefused(t *testing.T) { for _, tc := range refusedZeros { t.Run(tc.key, func(t *testing.T) { t.Parallel() - _, err := Load(writeYAML(t, yamlAt(t, tc.key, tc.zero))) + _, err := Load(writeYAML(t, yamlAt(t, tc.key, tc.zero, tc.also))) require.ErrorContains(t, err, tc.err, "the zero reaches Validate instead of becoming the default") }) } @@ -296,6 +315,8 @@ func parseDocDefault(t *testing.T, key, cell string, like any) any { v, err = strconv.ParseInt(cell, 10, 64) case float64: v, err = strconv.ParseFloat(cell, 64) + case time.Duration: + v, err = time.ParseDuration(cell) default: rt := reflect.TypeOf(like) switch { From c56103e90dda639cf182bca41f84b3d8b3741998 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:56:05 -0400 Subject: [PATCH 069/108] docs(dedupe): scope the DynamoDB Reserve undo to what it guarantees The comments on Reserve's undo, and the architecture page, said every put that may have landed is released. That holds for a sibling's failure, which never cuts off a put already sent, but not for a put cut off by the caller's cancellation or its own call deadline: DynamoDB can still apply it after the undo's conditional delete, and it then holds its key InFlight until the lease ends. A release that fails leaves a claim the same way. Both are bounded by the lease, as a crashed request's claims are, which the Deduplicator contract allows; the comments now say so instead of promising more. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 17 ++++++++++------- internal/dedupe/dynamodb_test.go | 2 +- 3 files changed, 12 insertions(+), 9 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index b8ab38cc..3c7330a6 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index f5c22b59..2181a1f5 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -320,8 +320,9 @@ type dynamoStore struct { // Reserve puts every key's pending item in parallel, each conditional on no // live item holding the key. A failed condition hands back the live item, // whose state says Duplicate or InFlight without a read, or Claimed when its -// token is the put's own. On any error every put that may have landed is -// released by its token. +// token is the put's own. On any error it releases, by token, every put that +// may have landed; what that undo misses (below) holds its key InFlight +// until the lease ends, as a crashed request's claim does. func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { if len(keys) == 0 { return []Claim{}, nil @@ -338,9 +339,10 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati sent := make([]bool, len(keys)) // The first failure skips the puts not yet sent: the Reserve fails // either way, and a throttled table should not take the rest. A put - // already sent runs on ctx, not gctx, so it finishes and its outcome is - // known before the undo below; cancelled mid-flight, it could land after - // its release and hold the key for the lease. + // already sent runs on ctx, not gctx, so a sibling's failure never cuts + // it off: it answers before the undo below. The caller's cancellation or + // the put's own deadline can, and DynamoDB may then apply it after its + // release. g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { @@ -363,8 +365,9 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati }) } if err := g.Wait(); err != nil { - // A put that was sent and errored may still have landed; its token - // is known, and releasing a key it does not hold is a no-op. + // A sent put that errored may have landed anyway; releasing a key + // its token does not hold is a no-op. Best effort: a put applied + // after this, or a release that fails, lapses with the lease. var undo []Claim for i, c := range claims { if sent[i] && (c.Status == Claimed || c.Status == 0) { diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 31f54f85..1a30791b 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -681,7 +681,7 @@ func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { // One throttled put in a multi-key Reserve: the unsent puts are never sent, // and a sibling already sent runs to its answer before the undo releases it, -// so a put cannot land after its own release. +// so a sibling's failure never lets a put land after its own release. func TestDynamo_FailedMultiKeyReserve(t *testing.T) { t.Parallel() var mu sync.Mutex From 8b616b4a309d7c7e96f964817d8f57c30d8cf5e1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:57:19 -0400 Subject: [PATCH 070/108] fix(config): cap dedupe.lease at 59.5s over the embedded queue The cap sat at the duplicate window itself (2m), which is the wrong boundary. The in-flight 503 answers with the whole lease as Retry-After, so a client that obeys it after a publish whose outcome it never learned republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second. With mq.backend=embedded a lease is now refused unless twice it plus one second fits the 2m window. The 30s default is unchanged; 59s and 59.5s pass, 60s is refused. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 4 ++-- docs/src/content/docs/sdk/reference.md | 2 +- internal/config/backends.go | 15 ++++++++++----- internal/config/backends_test.go | 7 +++++-- 7 files changed, 21 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index db2333b6..2abb1c9b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/config.yaml b/config.yaml index 7107dd38..0ca27bd6 100644 --- a/config.yaml +++ b/config.yaml @@ -57,7 +57,7 @@ mq: backend: embedded # NATS JetStream under /nats dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) - lease: 30s # how long a claimed id stays pending; at most 2m with the embedded mq + lease: 30s # how long a claimed id stays pending; at most 59.5s with the embedded mq (2*lease + 1s within its 2m duplicate window) reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend; no effect yet (ingest sends one id per call) # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 1484f79d..934ee671 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -120,7 +120,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: `dedupe.lease` may not exceed the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`) while `mq.backend` is `embedded`. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: while `mq.backend` is `embedded`, twice `dedupe.lease` plus one second must fit the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`), a cap of 59.5s (`maxEmbeddedLease`), because a client obeying the in-flight `503`'s `Retry-After` republishes up to twice the lease after the claim and DynamoDB rounds a claim's expiry up to the second. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 7dbd3171..12c0eeb0 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -56,7 +56,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` refuses boot. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, twice the lease plus one second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59.5s`: a client that obeys `Retry-After` after a publish whose outcome it never learned republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | | `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe @@ -265,7 +265,7 @@ cache: dedupe: backend: pebble # in-process Pebble under /pebble; or dynamodb - lease: 30s # at most 2m with the embedded mq + lease: 30s # at most 59.5s with the embedded mq reserve_concurrency: 64 # dynamodb: # read only when backend is dynamodb # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index 5679d785..390b0411 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -40,7 +40,7 @@ The SDK **never throws** for anything the server returns — all API errors come | 502 | `clickhouse.misconfigured` | No | ClickHouse refused WaveHouse's own credentials or database, or the route to it is wrong (a redirect, or a `4xx` other than `408`/`413`/`429`, with no exception code) — an operator fix | | 502 | `clickhouse.response_too_large` | No | A raw-SQL (`wh.sql`) response over the 64 MiB cap | | 503 | `clickhouse.unavailable` | Yes | ClickHouse is down, unreachable or overloaded; `Retry-After: 5`, honored between attempts | -| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | +| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the server's dedupe lease, 30 s by default). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits that long; a stream re-dials on its own jittered backoff instead | | 0 | `NETWORK_ERROR` | Yes | Network failure (retried with exponential backoff) | | 0 | `ABORTED` | No | Request canceled via `AbortSignal` | | 0 | `SSE_CONNECT_ERROR` | No | Stream could not be started (e.g. a non-absolute `baseURL`) | diff --git a/internal/config/backends.go b/internal/config/backends.go index 493e9488..5c1cad8b 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -165,11 +165,16 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) } -// embeddedDuplicateWindow mirrors mq.EmbeddedDuplicateWindow, the embedded -// ingest stream's duplicate window (#613 F2). A lease longer than it would let -// the republish of a publish whose outcome was unknown land twice. +// embeddedDuplicateWindow is the embedded ingest stream's duplicate window, +// counted from the stored publish. const embeddedDuplicateWindow = 2 * time.Minute +// maxEmbeddedLease is the longest dedupe.lease that window covers: a client +// that obeys the in-flight 503's Retry-After (the whole lease) after a +// publish whose outcome it never learned republishes up to twice the lease +// after the claim, and DynamoDB rounds a claim's expiry up to the second. +const maxEmbeddedLease = (embeddedDuplicateWindow - time.Second) / 2 + // validateBackends checks every layer's backend and its sub-block, then the // rules that span two layers. func (c *Config) validateBackends() error { @@ -178,8 +183,8 @@ func (c *Config) validateBackends() error { return err } } - if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > embeddedDuplicateWindow { - return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s exceeds the embedded mq's %s duplicate window: a claim must lapse before the queue forgets the publish it guards", c.Dedupe.Lease, embeddedDuplicateWindow) + if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > maxEmbeddedLease { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s is over %s with the embedded mq: twice the lease plus 1s must fit its %s duplicate window, since a client obeying the in-flight 503's Retry-After republishes up to twice the lease after the claim", c.Dedupe.Lease, maxEmbeddedLease, embeddedDuplicateWindow) } return nil } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 56f0d083..f4c01b0c 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -259,7 +259,8 @@ func TestValidate_Dedupe(t *testing.T) { c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true }, ""}, {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB = DedupeDynamoDBConfig{CreateTable: true} }, ""}, - {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, ""}, + {"lease just under a minute", func(c *Config) { c.Dedupe.Lease = 59 * time.Second }, ""}, + {"lease at the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond }, ""}, {"create_table without an endpoint", func(c *Config) { dynamo(c) c.Dedupe.DynamoDB.CreateTable = true @@ -275,7 +276,9 @@ func TestValidate_Dedupe(t *testing.T) { {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0"}, {"zero concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = 0 }, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0"}, {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, - {"lease past the duplicate window", func(c *Config) { c.Dedupe.Lease = 3 * time.Minute }, "exceeds the embedded mq's 2m0s duplicate window"}, + {"lease of a minute", func(c *Config) { c.Dedupe.Lease = time.Minute }, "dedupe.lease (WH_DEDUPE_LEASE) 1m0s is over 59.5s with the embedded mq: twice the lease plus 1s must fit its 2m0s duplicate window"}, + {"lease just past the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond + 1 }, "is over 59.5s with the embedded mq"}, + {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, "is over 59.5s with the embedded mq"}, } for _, tc := range cases { t.Run(tc.name, func(t *testing.T) { From 96b99683ba98b88f48ecd3fa333471551bedd75a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:01:23 -0400 Subject: [PATCH 071/108] docs(dedupe): the DynamoDB key example keeps '-' now keyenc keeps '-' since the merge from the base, so the example id evt-123 is stored as acme/clicks/evt-123, not acme/clicks/evt%2D123. Fixed in the Deployment page's table of attributes, which now also says what the escaping keeps, and in the CHANGELOG entry. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/deployment.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 83a34844..2be43f45 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 913ec294..cff048d6 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -453,7 +453,7 @@ What the backend requires of the table: | Attribute | Type | Role | |---|---|---| -| `pk` | String | Partition key, and the only key: tenant, table and id as readable text, for example `acme/clicks/evt%2D123` (the table and id escaped the way NATS subject tokens are). No sort key. | +| `pk` | String | Partition key, and the only key: tenant, table and id as readable text, for example `acme/clicks/evt-123` (the table and id escaped the way NATS subject tokens are: letters, digits, `_` and `-` kept, every other byte written as `%XX`). No sort key. | | `st` | Number | `1` = pending claim, `2` = committed. | | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | From a6944186d00d29b1fbe9fc7e8a39c07a5a1d2238 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:07:53 -0400 Subject: [PATCH 072/108] fix(app): the dedupe reload hook makes no DynamoDB call While the table check failed, every settings reload ran it from the AfterAdopt hook, which the registry runs under the lock that serializes reloads: 2.5s per reload against a hung endpoint at the default timeout, the hooks registered after it waiting, and a tenant the reload switched on answering ErrDisabled meanwhile, so its records published un-deduped. The hook now applies every store against the last check's result, which fails a switched-on store closed (ErrUnavailable) while the check has not passed, and wakes the background retry through a one-slot channel, so a reload still retries at once. Only boot and that loop run the check, and once passed it stays passed. A flat directory's boot is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/dedupe_dynamodb_test.go | 82 ++++++++++++++++---- internal/app/wire.go | 70 ++++++++++------- 7 files changed, 118 insertions(+), 44 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2abb1c9b..d980aab2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload never waits on the table: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 934ee671..0d0d7741 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -93,7 +93,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes: by a background component that backs off from one second to thirty (a nested directory has no watcher), and by every reload. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads: it applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 12c0eeb0..2fd31f25 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -61,7 +61,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background (backing off from one second to thirty) and on every reload, so a table that comes good is picked up without a restart. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload never waits on the table: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index cf617b0d..f2100f4c 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload, which never waits on the table). No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index dcba15d4..b6f5397e 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index 6bac9dc2..f3fb1e59 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -23,10 +23,11 @@ import ( // fakeDynamo answers the DynamoDB JSON protocol for one table, enough for // boot's check, the dev create path, and a claim and its commit. Whether the -// table exists is the test's to switch. +// table exists, and whether the endpoint hangs, are the test's to switch. type fakeDynamo struct { mu sync.Mutex exists bool + hangs bool calls []string } @@ -36,15 +37,24 @@ func (f *fakeDynamo) setExists(v bool) { f.exists = v } -func (f *fakeDynamo) called(op string) bool { +func (f *fakeDynamo) setHangs(v bool) { f.mu.Lock() defer f.mu.Unlock() + f.hangs = v +} + +func (f *fakeDynamo) called(op string) bool { return f.count(op) > 0 } + +func (f *fakeDynamo) count(op string) int { + f.mu.Lock() + defer f.mu.Unlock() + n := 0 for _, c := range f.calls { if c == op { - return true + n++ } } - return false + return n } func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { @@ -55,8 +65,12 @@ func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { if op == "CreateTable" { f.exists = true } - exists := f.exists + exists, hangs := f.exists, f.hangs f.mu.Unlock() + if hangs { + <-r.Context().Done() + return + } w.Header().Set("Content-Type", "application/x-amz-json-1.0") if !exists { w.WriteHeader(http.StatusBadRequest) @@ -144,10 +158,10 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { }) } }) - t.Run("nested fails closed until a reload passes the check", func(t *testing.T) { + t.Run("nested fails closed", func(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil}) cfg := testConfig(t, root) - fake := dynamoConfig(t, cfg, false) + dynamoConfig(t, cfg, false) a := newApp(t, cfg, Options{}) acme := a.dedup.For("acme") @@ -156,15 +170,57 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on, table missing: ingest fails closed") _, err = dedupetest.Mark(t.Context(), a.dedup.For("globex"), eventKey) require.ErrorIs(t, err, dedupe.ErrDisabled) - - fake.setExists(true) - a.tenants.Reload("test") - assert.True(t, acme.Open(), "the reload checked again and opened the store") - _, err = dedupetest.Mark(context.Background(), acme, eventKey) - require.NoError(t, err) }) } +// The reload hook runs under the lock that serializes reloads, so it never +// calls DynamoDB: against a table that hangs, a reload returns at once, and a +// tenant it switches on fails closed rather than publishing un-deduped. +func TestReload_DynamoDBDedupeMakesNoTableCall(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + a := newApp(t, cfg, Options{}) + fake.setHangs(true) + before := fake.count("DescribeTable") + + rewriteSettings(t, filepath.Join(root, "acme"), dedupeOn) + start := time.Now() + a.tenants.Reload("test") + assert.Less(t, time.Since(start), time.Second, "a check would wait out its 2.5s deadline") + assert.Equal(t, before, fake.count("DescribeTable"), "the reload made no table call") + + acme := a.dedup.For("acme") + assert.False(t, acme.Open()) + _, err := dedupetest.Mark(t.Context(), acme, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on while the table fails: closed, not ErrDisabled") +} + +// A reload wakes the background retry rather than running the check itself. +// The retry's first timed attempt is a second after it starts, and a timer +// never fires early, so an open sooner than that is the reload's doing. +func TestRun_DynamoDBDedupeReloadWakesTheRetry(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + acme := a.dedup.For("acme") + require.False(t, acme.Open()) + + start := time.Now() + _, stop := runApp(t, a, ln) + fake.setExists(true) + a.tenants.Reload("test") + require.Eventually(t, acme.Open, 5*time.Second, 5*time.Millisecond) + assert.Less(t, time.Since(start), time.Second, "opened before the first timed retry") + _, err = dedupetest.Mark(context.Background(), acme, eventKey) + require.NoError(t, err) + require.NoError(t, stop()) +} + // A nested directory has no watcher, so a table that comes good is picked up // by the background retry, not only by a reload someone has to send. func TestRun_DynamoDBDedupeRetriesTheTableCheck(t *testing.T) { diff --git a/internal/app/wire.go b/internal/app/wire.go index 089de185..08dd3763 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -545,7 +545,10 @@ var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") // store closed, so its ingest fails closed. Unlike a local disk, a remote // table's failure is usually brief (a throttle, credentials not yet issued // mid-rollout), and a nested directory has no watcher to reload it, so the -// check is also retried in the background, with backoff, until it passes. +// check is then retried in the background, with backoff, until it passes. +// The check is network I/O, so the AfterAdopt hook never runs it: the hook +// holds the lock that serializes reloads. It applies every store against the +// last check's result and wakes the retry, so a reload still retries at once. func (a *App) wireDynamoDedupe(ctx context.Context) error { c := a.cfg.Dedupe.DynamoDB d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ @@ -557,72 +560,87 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { return err } var mu sync.Mutex - state := errDynamoUnchecked // nil once the table has passed - check := func(ctx context.Context) error { + state := errDynamoUnchecked // nil once the table has passed, for good + ready := func() error { mu.Lock() defer mu.Unlock() - if state == nil { - return nil - } - if c.CreateTable { - if state = d.CreateTable(ctx); state != nil { - return state - } - } - state = d.Check(ctx) return state } - ready := func() error { + // check is only ever run by boot, then by the retry loop, one at a time. + check := func(ctx context.Context) error { + var err error + if c.CreateTable { + err = d.CreateTable(ctx) + } + if err == nil { + err = d.Check(ctx) + } mu.Lock() defer mu.Unlock() + if state != nil { + state = err + } return state } stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) a.dedup = stores a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) - var reconciling sync.Mutex // the hook and the retry loop both reconcile - reconcile := func(ctx context.Context) error { + var reconciling sync.Mutex // the hook and the retry loop both apply + apply := func() { reconciling.Lock() defer reconciling.Unlock() if err := stores.Retain(a.served); err != nil { slog.Error("dedupe store close failed", "error", err) } - checkErr := check(ctx) - if checkErr != nil { - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", - "table", c.Table, "error", checkErr) - } for id, store := range a.tenants.All() { m := stores.For(id) enabled := store.DedupeEnabled() wasOpen := m.Open() - // The one failure an open has is the check's, logged above. + // The one failure an open has is the check's, logged where it ran. _ = m.Apply(enabled) if m.Open() != wasOpen { slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) } } - return checkErr } - a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) - if err := reconcile(ctx); err != nil { + retry := make(chan struct{}, 1) + a.tenants.AfterAdopt(func([]tenant.ID) { + apply() + if ready() != nil { + select { + case retry <- struct{}{}: + default: // a retry is already due + } + } + }) + if err := check(ctx); err != nil { if !a.tenants.Nested() { return fmt.Errorf("dedupe open: %w", err) } + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + "table", c.Table, "error", err) a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { select { case <-ctx.Done(): return nil case <-time.After(wait): + case <-retry: } - if reconcile(ctx) == nil { - slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + if err := check(ctx); err != nil { + if ctx.Err() == nil { + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + "table", c.Table, "error", err) + } + continue } + slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + apply() } return nil }}) } + apply() return nil } From 9b8ea3c5b16693218acfe7984bae9a3b517809e9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:08:25 -0400 Subject: [PATCH 073/108] fix(app): say what a failed dedupe table check leads to The log said ingest with dedupe on fails closed until a reload passes the check, which is wrong in both shapes: a flat directory refuses boot, and a nested one retries the check in the background. The boot line and the retry's line now say that it is retried and ingest fails closed meanwhile. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/app/wire.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/internal/app/wire.go b/internal/app/wire.go index 08dd3763..a65d1bf8 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -617,7 +617,7 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { if !a.tenants.Nested() { return fmt.Errorf("dedupe open: %w", err) } - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", "table", c.Table, "error", err) a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { @@ -629,7 +629,7 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { } if err := check(ctx); err != nil { if ctx.Err() == nil { - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + slog.Error("dedupe: dynamodb table check failed again; ingest with dedupe on still fails closed", "table", c.Table, "error", err) } continue From 0425a555f9e03bc281c8f507acdce388846d03ae Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:17:32 -0400 Subject: [PATCH 074/108] fix(dedupe): never size the DynamoDB idle pool below the SDK default MaxIdleConnsPerHost was set to ReserveConcurrency outright, so a small ReserveConcurrency (a boot key soon) shrank the pool below the SDK's default of 10 per host. It is now the larger of the two, as MaxIdleConns already was. TestNewDynamo_SizesTheIdlePool adds 4 and asserts both limits stay at or above the SDK's defaults; it failed at 4 and 8 before. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 9 +++++---- internal/dedupe/dynamodb_test.go | 4 +++- 3 files changed, 9 insertions(+), 6 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index cbc0d55c..7935f25c 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -137,7 +137,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt-123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped and joined (`keyenc.AppendJoin`) by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with at least as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 2181a1f5..cc3321f7 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -81,7 +81,8 @@ type DynamoConfig struct { // the client after throttles. RetryMode string // ReserveConcurrency bounds the parallel calls one Reserve, Commit or - // Release makes, and sizes the client's idle connection pool to match. + // Release makes, and sizes the client's idle connection pool to match + // (never below the SDK's default of 10 per host). // 0 = 64. ReserveConcurrency int } @@ -158,11 +159,11 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load } // newHTTPClient keeps an idle connection for every call one Reserve can have -// in flight: with the SDK's default of 10 per host, a wide Reserve would dial -// most of its puts afresh. +// in flight, and never fewer than the SDK's defaults: with its 10 per host, a +// wide Reserve would dial most of its puts afresh. func newHTTPClient(cfg DynamoConfig) *awshttp.BuildableClient { return awshttp.NewBuildableClient().WithTransportOptions(func(tr *http.Transport) { - tr.MaxIdleConnsPerHost = cfg.ReserveConcurrency + tr.MaxIdleConnsPerHost = max(tr.MaxIdleConnsPerHost, cfg.ReserveConcurrency) tr.MaxIdleConns = max(tr.MaxIdleConns, cfg.ReserveConcurrency) }) } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 1a30791b..03a7196f 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -585,7 +585,7 @@ func TestDynamo_ThrottledCallEndsOnItsLastAttempt(t *testing.T) { // extra (TestDynamo_ThrottledCallEndsOnItsLastAttempt's) replaces it. func TestNewDynamo_SizesTheIdlePool(t *testing.T) { t.Parallel() - for _, n := range []int{0, 8, 200} { + for _, n := range []int{0, 4, 8, 200} { d, err := NewDynamo(t.Context(), DynamoConfig{Table: "dedupe", Region: "us-east-1", ReserveConcurrency: n}) require.NoError(t, err) client, ok := d.api.(*dynamodb.Client).Options().HTTPClient.(*awshttp.BuildableClient) @@ -593,6 +593,8 @@ func TestNewDynamo_SizesTheIdlePool(t *testing.T) { tr := client.GetTransport() assert.GreaterOrEqual(t, tr.MaxIdleConnsPerHost, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) assert.GreaterOrEqual(t, tr.MaxIdleConns, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) + assert.GreaterOrEqual(t, tr.MaxIdleConnsPerHost, awshttp.DefaultHTTPTransportMaxIdleConnsPerHost, "never below the SDK's default: ReserveConcurrency %d", n) + assert.GreaterOrEqual(t, tr.MaxIdleConns, awshttp.DefaultHTTPTransportMaxIdleConns, "ReserveConcurrency %d", n) } } From 5c64ace0fa3f47e121d4103d4c62d2d28b0365b5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:17:55 -0400 Subject: [PATCH 075/108] docs(dedupe): list the short-circuit counter with the DynamoDB metrics The Metrics bullet on the Deployment page named three of the backend's four metrics; wavehouse_dedupe_dynamodb_short_circuits_total appeared only in passing one bullet earlier. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/deployment.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index cff048d6..e2be2616 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -510,7 +510,7 @@ data "aws_iam_policy_document" "wavehouse_dedupe" { - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. - **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. - **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five throttled or unreachable claims in a row within one second, the backend stops calling the table for a second and fails every tenant's dedupe requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). A duplicate or in-flight answer is not a failure and resets the count. -- **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. +- **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`, `wavehouse_dedupe_dynamodb_short_circuits_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. ## Upgrading across the v2 ingest envelope From 860d92985818df094541427f876a69ecbd48477b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:20:22 -0400 Subject: [PATCH 076/108] test(mq): state the lease/duplicate-window invariant as 2*lease+1s The in-flight 503 for an uncertain publish sends the FULL dedupe lease as Retry-After, so an obedient client's retry can land up to ~2*lease after the original Reserve -- not just one lease later -- and a claim's expiry can itself round up by up to a second on some backends (DynamoDB, for one). "lease <= window" understates what the embedded queue's duplicate window actually has to cover. Pin the real invariant in TestIngest_DedupeLeaseFitsTheDuplicateWindow (2*dedupe.DefaultLease + time.Second <= mq.EmbeddedDuplicateWindow), correct the embedded.go comment that said "must not exceed", and reword the same claim in durability.md, api.md, architecture.md and CHANGELOG.md. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 ++-- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 2 +- internal/api/ingest_window_test.go | 12 ++++++++---- internal/mq/embedded.go | 28 ++++++++++++++++++++------- 6 files changed, 34 insertions(+), 16 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index dac04cfc..2fddb0e3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -86,7 +86,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md#upgrading-across-the-dedupe-key-change)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_ingest_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_ingest_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, sized to `2 × the 30-second lease + 1s`: an uncertain publish's `503` sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and the `+1s` covers a backend whose claim expiry itself rounds up by that much. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_ingest_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index efdd4373..fd8d37c0 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -295,10 +295,10 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store is not open (for example, it failed to open on a reload); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue remembers the key for two minutes after the first publish, so a retry inside that window stores no second copy (the SDK's, after the 30-second `Retry-After`, lands inside it); a later one is stored again. | +| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown, other than a full queue or an unreachable broker (below): the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue's duplicate window (two minutes) covers up to ~2×lease plus a margin, not just the lease itself, so a retry timed off `Retry-After` anywhere in this flow stores no second copy; a much later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | -| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time (a transient broker failure, not a full queue). Response includes `Retry-After: 5` header. Reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. With dedupe on, the record's id is given back so the retry can publish — but an unavailable broker that timed out may already have stored the event, so that retry can publish a second copy (the windowed-ingest follow-up closes this with an idempotency key). | +| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time (`mq.ErrUnavailable`, a transient broker failure, not a full queue) — reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. As for the `500`, the record's id is left to lapse rather than given back, so a retry cannot land as a second copy; `Retry-After` is that lease, rounded up to whole seconds, when dedupe was on for the record, else the flat `Retry-After: 5`. | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 15ab9574..906469bd 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -164,7 +164,7 @@ The **only** package that imports NATS/JetStream — a `depguard` rule in `.gola - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: ASCII letters, digits, `_` and `-` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **deadletter.go** — `deadLetterTables`, the per-table count `DeadLetterCounts` reports: a dead-letter stream's per-subject counts, each subject parsed back to its topic and counted under its table — every scope of a table under the table itself, so a dotted table name never shares a count with a table + scope pair — and a table filter keeps that table with all of its scopes. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`) and remembering idempotency keys for `EmbeddedDuplicateWindow` (two minutes, which a dedupe lease must not exceed), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`) and remembering idempotency keys for `EmbeddedDuplicateWindow` (two minutes, sized to `2 × the dedupe lease + 1s` — the in-flight `503` sends the full lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original `Reserve`, and the `+1s` covers a backend whose claim expiry itself rounds up by that much), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. - **mqtest/** — The conformance suite for `Broker` (`mqtest.Run`): the behavior the rest of the process relies on — publish and consume round trips with names that need encoding, per-tenant order, redelivery, dead-lettering and its counts, replay bounds and isolation, the one `failed` report of a consumer whose delivery ends underneath it — checked through the interfaces alone, with no stream or subject name in sight. Each implementation runs it from a test of its own — the embedded one from `mqtest/embedded_test.go`, a test binary apart from `internal/mq`'s so the two share no 15s budget — handing it a fresh broker per case and flags (`mqtest.Caps`) for the few places where backends legitimately differ: whether a full queue refuses its own tenant alone, whether `PurgeAcked` removes anything, whether a tenant never given a budget has a dead-letter queue to report on, and whether `CreateConsumer` configures the durable or only finds one. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index e6a28c58..324e1db1 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_ingest_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The duplicate window has to cover more than the lease alone: the `503` for an uncertain publish sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and a claim's expiry can itself round up by a further second on some backends — the invariant the queue configuration and its tests pin is `2×lease + 1s ≤ window`, not just `lease ≤ window`. Two minutes against a 30-second lease clears that with room to spare. ## Check your storage before you trust it diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index 3d1dc931..ba350eed 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -249,12 +249,16 @@ func TestIngest_Windows_OutcomesStayInOrder(t *testing.T) { } } -// The embedded queue must remember an idempotency key for at least a lease: -// the retry of an uncertain publish lands after the lease, and only the queue's -// duplicate window drops its second copy. +// The embedded queue must remember an idempotency key for at least two +// leases plus a second: the in-flight 503 of an uncertain publish sends the +// full lease as Retry-After, so a client that obeys it can republish up to +// ~2*lease after the original Reserve, and a claim's expiry can itself round +// up by up to a second (a DynamoDB backend, for one). Only the queue's +// duplicate window running at least that long guarantees it still drops the +// retry's second copy. func TestIngest_DedupeLeaseFitsTheDuplicateWindow(t *testing.T) { t.Parallel() - assert.LessOrEqual(t, dedupe.DefaultLease, mq.EmbeddedDuplicateWindow) + assert.LessOrEqual(t, 2*dedupe.DefaultLease+time.Second, mq.EmbeddedDuplicateWindow) } // faultyPublisher publishes through a real broker and fails the calls fail diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 304e1f13..1e3853ed 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -227,6 +227,13 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { held uint64 } dlqs := map[tenant.ID]dlqState{} + // duplicates is the ingest stream's own Duplicates window as found on + // disk, keyed alongside dlqs: a stream from before EmbeddedDuplicateWindow + // existed, or reopened under a different value, must not be counted as + // already at budget below, or SetMaxBytes(same budget) short-circuits and + // the stale window is never brought forward (measured: a stream with + // Duplicates=10s kept 10s after NewEmbedded + SetMaxBytes(same budget)). + duplicates := map[tenant.ID]time.Duration{} streams := e.js.ListStreams(ctx) for info := range streams.Info() { name := info.Config.Name @@ -234,6 +241,7 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { q := e.queue(id) q.ingest = true q.asked, q.ingestCap = info.Config.MaxBytes, info.Config.MaxBytes + duplicates[id] = info.Config.Duplicates } else if id, ok := streamTenant(dlqStreamPrefix, name); ok { e.queue(id).dlq = true dlqs[id] = dlqState{limit: info.Config.MaxBytes, held: info.State.Bytes} @@ -244,14 +252,16 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { } // A pair is at its budget when its dead-letter stream is at a tenth of // the ingest cap, or above it holding more than that: the shrink guard's - // doing. Anything else is a pair a stop or a failed update left split, or - // one missing its dead-letter stream, so its budget stays unapplied and - // the boot's SetMaxBytes applies it to both streams again. + // doing, AND its ingest stream's duplicate window already matches + // EmbeddedDuplicateWindow. Anything else is a pair a stop or a failed + // update left split, one missing its dead-letter stream, or one whose + // duplicate window is stale, so its budget stays unapplied and the boot's + // SetMaxBytes applies it — and the current window — to both streams again. for id, q := range e.queues { d, ok := dlqs[id] tenth := q.asked / dlqShare guarded := d.limit > tenth && d.held <= math.MaxInt64 && int64(d.held) > tenth - if q.ingest && ok && (d.limit == tenth || guarded) { + if q.ingest && ok && (d.limit == tenth || guarded) && duplicates[id] == EmbeddedDuplicateWindow { q.maxBytes = q.asked } e.record(id, q) @@ -318,9 +328,13 @@ func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { } // EmbeddedDuplicateWindow is how long an ingest queue remembers a -// WithIdempotencyKey key. A dedupe lease must not exceed it: a claim left to -// lapse after an uncertain publish is republished once the lease ends, and -// only this window drops that second copy. +// WithIdempotencyKey key. It must be at least 2*lease + 1s: a claim left to +// lapse after an uncertain publish is republished once the lease ends, but +// the in-flight 503 tells a client to retry only after the FULL lease, so an +// obedient client's retry can land up to ~2*lease after the original +// Reserve; the +1s covers a backend (DynamoDB, for one) that rounds a +// claim's expiry up by as much. Only a window at least that long guarantees +// this queue still drops the retry's second copy. const EmbeddedDuplicateWindow = 2 * time.Minute // ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard From 65524789d0c9d3f836b56df5859a130219d3da49 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:19 -0400 Subject: [PATCH 077/108] test(mq): pin the takeStock boot fix for a stale duplicate window The takeStock fix landed with the previous commit (embedded.go), since both are about the same staleness question; this adds its coverage. TestNewEmbedded_TakeStockRefreshesAStaleDuplicateWindow reopens a store whose ingest stream was left with a Duplicates window other than EmbeddedDuplicateWindow, then calls SetMaxBytes with the SAME budget as before and checks the window is brought forward -- the path TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat did not cover (its stream is created directly, never recorded by takeStock, so SetMaxBytes's same-budget early return never applies to it). Corrected that test's comment to say so, pointing at the new one. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/mq/embedded_test.go | 36 ++++++++++++++++++++++++++++++++++-- 1 file changed, 34 insertions(+), 2 deletions(-) diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index cf96ee9d..7c7a87a0 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -163,8 +163,12 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { } // A repeated idempotency key inside the duplicate window is dropped as a -// success, so an uncertain publish can be republished safely; a queue made -// with another window gets this one on its next budget apply. +// success, so an uncertain publish can be republished safely. This stream is +// created directly, never recorded by takeStock, so SetMaxBytes's next +// budget apply always runs and picks up the current window; +// TestNewEmbedded_TakeStockRefreshesAStaleDuplicateWindow covers the boot +// path, where takeStock itself must not mistake a stale window for one +// already at budget. func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { e := openEmbedded(t, t.TempDir()) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) @@ -1500,6 +1504,34 @@ func TestNewEmbedded_TakesStockOfTheQueuesOnDisk(t *testing.T) { assert.Equal(t, int64(8<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) } +// takeStock must not count a stream as at its budget when its Duplicates +// window is stale (from before EmbeddedDuplicateWindow existed, or changed +// underneath it): otherwise SetMaxBytes's same-budget early return never lets +// a later apply bring the window forward, and the stream keeps whatever it +// had indefinitely. +func TestNewEmbedded_TakeStockRefreshesAStaleDuplicateWindow(t *testing.T) { + t.Parallel() + dir := storeDir(t) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + first, err := NewEmbedded(dir) + require.NoError(t, err) + require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) + stale := ingestStreamConfig("acme", 8<<20) + stale.Duplicates = 10 * time.Second + _, err = first.js.UpdateStream(ctx, stale) + require.NoError(t, err) + require.NoError(t, first.Close()) + + e := openEmbedded(t, dir) + require.Equal(t, 10*time.Second, streamConfig(t, e, "INGEST_acme").Duplicates, "the stale window is still on disk") + + require.NoError(t, e.SetMaxBytes(ctx, "acme", 8<<20), "same budget as before") + assert.Equal(t, EmbeddedDuplicateWindow, streamConfig(t, e, "INGEST_acme").Duplicates, + "takeStock must not have marked this pair already at budget, or this apply would have no-op'd") +} + // A durable found on disk is kept as it stands when it holds the settings // asked for — a boot over many queues writes nothing it need not — and is // updated in place when they differ; either way delivery resumes past what it From 67edd57b698590d660a24269354f865a1c5484da Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:29 -0400 Subject: [PATCH 078/108] test(mq): add an idempotency-key case to the Broker conformance suite mqtest.Run had no case for mq.WithIdempotencyKey, yet the uncertain- publish rule in the ingest handler (a claim left to lapse after a publish whose outcome is unknown, relying on the queue to drop the retry's second copy) depends on every Broker honouring it -- not just the embedded one, which already had its own duplicate-key test. Adds IdempotencyKeyDropsARepeat: a publish repeated under one key is stored once and both calls return nil; a different key is stored separately, checked via ReplaySince. Wired into the suite's case list unconditionally (no Caps flag), since every Broker must honour it. Kept internal/mq's own TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat: it additionally pins the embedded-specific mechanics of a stream opened with one duplicate window picking up EmbeddedDuplicateWindow on its next SetMaxBytes, which is outside the generic Broker contract. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/mq/mqtest/cases.go | 17 +++++++++++++++++ internal/mq/mqtest/mqtest.go | 1 + 2 files changed, 18 insertions(+) diff --git a/internal/mq/mqtest/cases.go b/internal/mq/mqtest/cases.go index 699b5194..513501de 100644 --- a/internal/mq/mqtest/cases.go +++ b/internal/mq/mqtest/cases.go @@ -157,6 +157,23 @@ func roundTrip(t *testing.T, h Harness) { } } +// A publish repeated with the same idempotency key inside the duplicate +// window is a no-op reported as success: the ingest handler relies on this to +// make a retry of an uncertain publish (the outcome unknown after a failure +// other than a full queue) safe rather than a second copy. A different key +// is its own event. +func idempotencyKeyDropsARepeat(t *testing.T, h Harness) { + b := h.New(t) + topic := mq.Topic{Tenant: Acme, Table: "idem"} + + require.NoError(t, b.Publish(ctx(t), topic, []byte("first"), mq.WithIdempotencyKey("k1"))) + require.NoError(t, b.Publish(ctx(t), topic, []byte("repeat"), mq.WithIdempotencyKey("k1")), + "a repeat under the same key is reported as success, not stored again") + require.NoError(t, b.Publish(ctx(t), topic, []byte("second"), mq.WithIdempotencyKey("k2"))) + + replayEventually(t, b, topic, time.Time{}, []string{"first", "second"}) +} + // Nothing lands on a tenant by omission (#583), and an invalid tenant is not // backpressure a retry could clear. func refusesATopicWithoutATenant(t *testing.T, h Harness) { diff --git a/internal/mq/mqtest/mqtest.go b/internal/mq/mqtest/mqtest.go index 5978a5b6..1d9f2edd 100644 --- a/internal/mq/mqtest/mqtest.go +++ b/internal/mq/mqtest/mqtest.go @@ -83,6 +83,7 @@ func Run(t *testing.T, h Harness) { cases := []testCase{ {"RoundTrip", true, roundTrip}, + {"IdempotencyKeyDropsARepeat", true, idempotencyKeyDropsARepeat}, {"RefusesATopicWithoutATenant", true, refusesATopicWithoutATenant}, {"SubscribeCarriesTheTraceContext", true, subscribeCarriesTheTraceContext}, {"SubscribeSeesEveryTenant", true, subscribeSeesEveryTenant}, From a6b2108c527f1a06abcd1471b160c49f53af8f40 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:39 -0400 Subject: [PATCH 079/108] fix(app,api): correct a stale comment; don't ERROR-log a client-gone Reserve internal/app/wire.go's wirePebbleDedupe comment still described a store that fails to open on reload as a 500 "dedupe failed" -- that mapping moved to 503 "dedupe store unavailable" with Retry-After: 5 earlier in this branch's history (internal/api/reserve's dedupe.ErrUnavailable case). Update the comment to match. reserve()'s generic dd.Reserve error branch logged every failure at ERROR, including one caused by the request's own context ending (the client went away, or its deadline passed) while Reserve was in flight -- not a backend problem, and not worth paging an operator over. Check ctx.Err() and log at Debug instead when it is set; a real backend failure still logs ERROR. The response status is unchanged (moot: nothing is listening for it). TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError pins the log level via logtest; TestIngest_Dedup_ReserveError continues to pin the real-backend-failure ERROR case. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest.go | 11 ++++++++++- internal/api/ingest_test.go | 25 +++++++++++++++++++++++++ internal/app/wire.go | 5 +++-- 3 files changed, 38 insertions(+), 3 deletions(-) diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 9c92850c..2007bb7b 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -835,7 +835,16 @@ func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, tab slog.WarnContext(ctx, "dedupe store unavailable", "error", err, "table", table) return &requestAbort{Status: http.StatusServiceUnavailable, Message: "dedupe store unavailable", RetryAfter: "5"} case err != nil: - slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "table", table) + if ctx.Err() != nil { + // The request's own context ended — the client is gone, or its + // deadline passed — while Reserve was in flight. Reserve wraps + // that as an ordinary error, but it is not a backend problem + // worth an operator's attention, and the response status below + // is moot: nothing is listening for it. + slog.DebugContext(ctx, "dedupe reserve failed: request context ended", "error", err, "table", table) + } else { + slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "table", table) + } return &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} } var held *dedupe.Key diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index db206fd8..8d52815a 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -2968,6 +2968,31 @@ func TestIngest_Dedup_ReserveError(t *testing.T) { assert.Empty(t, pub.Published()) } +// A Reserve error caused by the request's own context ending (the client +// gone, or its deadline past) is not a backend failure and must not log at +// ERROR — an operator paging on ERROR logs would otherwise be woken by +// clients that simply went away. TestIngest_Dedup_ReserveError above pins the +// real-backend-failure case, which stays ERROR. +func TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError(t *testing.T) { + buf := logtest.Capture(t, slog.LevelDebug) + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.Err = errors.New("backend down") + h := dedupHandler(t, pub, dedup, false) + + ctx, cancel := context.WithCancel(context.Background()) + cancel() + req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}).WithContext(ctx) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(req)) + + assert.Contains(t, buf.String(), "dedupe reserve failed", "still logged, just not at ERROR") + assert.NotContains(t, buf.String(), `"level":"ERROR"`, "a client-gone Reserve error must not page an operator") + assert.Contains(t, buf.String(), `"level":"DEBUG"`) + assert.Empty(t, pub.Published()) +} + // #370: an explicit null id is a missing id — rejected under require_id, // published un-deduped otherwise — never the one id "" that made every // null record after the first a duplicate. diff --git a/internal/app/wire.go b/internal/app/wire.go index 20a9036f..506b875e 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -484,8 +484,9 @@ func (a *App) wireDedupe() error { // still closed — either the hook sees it or the boot apply reads it. An // instance that cannot open follows the registry's own rule for the shape: // flat refuses boot, like every other store, and on reload logs and leaves -// the store closed — ingest then fails closed (500 "dedupe failed") rather -// than silently publishing un-deduped, since the files asked for dedupe; +// the store closed — ingest then fails closed (503 "dedupe store +// unavailable", Retry-After: 5) rather than silently publishing un-deduped, +// since the files asked for dedupe; // nested fails closed the same way at boot too, for every tenant with // dedupe on, the next reload retrying, so it never costs the process. func (a *App) wirePebbleDedupe() error { From 04da8010494fef1a67f441bdafb675bc28753b20 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:41:29 -0400 Subject: [PATCH 080/108] test(dedupe): bound the commit-race case under low GOMAXPROCS The busy-loop workers had no start barrier, so at low GOMAXPROCS they starved the committer instead of racing it: the case alone took 2.1-2.8s at -cpu 1 (0.16s at -cpu 2), and the release-before-write mutant went from caught 6/6 at -cpu 4 to 0/6 at -cpu 1. Add a start barrier (each worker signals after its first Reserve returns; Commit waits for all of them) and drop worker/round counts (8x20 to 4x8), matching the measured fix. Also widen the pre-existing single-key race case to 20 fresh keys. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/dedupetest/dedupetest.go | 31 +++++++++++++----------- 1 file changed, 17 insertions(+), 14 deletions(-) diff --git a/internal/dedupe/dedupetest/dedupetest.go b/internal/dedupe/dedupetest/dedupetest.go index 5eca743e..94e0b27c 100644 --- a/internal/dedupe/dedupetest/dedupetest.go +++ b/internal/dedupe/dedupetest/dedupetest.go @@ -182,11 +182,8 @@ var cases = []struct { "retention 0 never expires") }}, {"concurrent reserves of one key claim it once", func(t *testing.T, s *suite) { - // #390: two requests carrying one id must not both publish. Run - // over many fresh keys, not just one: a race confined to a narrow - // lock window (e.g. one unlocked too early around a single - // backend read) can slip past a single key far more often than it - // triggers, so one key is a weak witness. + // #390: two requests carrying one id must not both publish, checked + // across many fresh keys since a narrow lock window can miss one. const n = 16 const keys = 20 d, p := s.store(t, "acme"), s.peer(t, "acme") @@ -226,21 +223,21 @@ var cases = []struct { } }}, {"a reserve racing a commit never claims, and settles to duplicate once it lands", func(t *testing.T, s *suite) { - // A backend's Commit must write the durable record before it drops - // the pending claim (e.g. the Pebble backend's synced batch flush - // ahead of releasing the in-memory claim) — a Reserve spinning - // against the same key during that window must never see Claimed, - // and must see Duplicate the instant Commit returns. Run over many - // fresh keys and both clients so a narrow unlock window isn't - // masked by luck on one key or one process's view. - const workers = 8 - const rounds = 20 + // Commit must land durably before it drops the pending claim; a + // racing Reserve must never see Claimed, only Duplicate once it + // returns. A start barrier holds Commit until every worker has + // made its first call, so low GOMAXPROCS can't starve them out of + // overlapping it at all. + const workers = 4 + const rounds = 8 d, p := s.store(t, "acme"), s.peer(t, "acme") for round := range rounds { k := key(fmt.Sprintf("commit-race-%d", round)) c := reserve(t, d, long, k) var stop atomic.Bool var claimed atomic.Int64 + var ready sync.WaitGroup + ready.Add(workers) var wg sync.WaitGroup for i := range workers { store := d @@ -248,8 +245,13 @@ var cases = []struct { store = p } wg.Go(func() { + first := true for !stop.Load() { got, err := store.Reserve(context.Background(), []dedupe.Key{k}, long) + if first { + first = false + ready.Done() + } if !assert.NoError(t, err) { return } @@ -263,6 +265,7 @@ var cases = []struct { } }) } + ready.Wait() require.NoError(t, d.Commit(t.Context(), c, 0)) stop.Store(true) wg.Wait() From f0f39ebfc18442a0cd99f0bab33e7d0c16d6a21a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:42:15 -0400 Subject: [PATCH 081/108] docs(api): the uncertain-publish caveat covers every publish failure MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The caveat about a released id letting a retry publish a genuine second copy sat only on the mq.ErrUnavailable 503 rows, which the same text says the embedded broker never returns — so as written it described a case the default deployment can't hit. With the embedded broker the uncertain publish is the 500 (a client disconnect after the broker had already stored the message), so state the caveat there too, cross- referenced from the 503 rows for the external-broker case. Link the follow-up as #629 instead of the unlinked "windowed-ingest follow-up". Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 8 ++++---- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9af2ab30..2ba49117 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,7 +85,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md#upgrading-across-the-dedupe-key-change)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_ingest_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)) — the residual case is a publish that fails after it already reached the broker (a timeout, a disconnect), where the released id lets the retry through but that retry publishes a genuine second copy; [#629](https://github.com/Wave-RF/WaveHouse/pull/629) closes that with an idempotency key. The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md#upgrading-across-the-dedupe-key-change)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_ingest_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index e63846b5..8afdaeb1 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -294,10 +294,10 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate. | +| 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate — but if the publish reached the broker before failing (e.g. a client disconnect after the embedded broker had already stored the message), that retry can publish a second copy ([#629](https://github.com/Wave-RF/WaveHouse/pull/629) closes this with an idempotency key). | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | -| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time (a transient broker failure, not a full queue). Response includes `Retry-After: 5` header. Reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. With dedupe on, the record's id is given back so the retry can publish — but an unavailable broker that timed out may already have stored the event, so that retry can publish a second copy (the windowed-ingest follow-up closes this with an idempotency key). | +| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time (a transient broker failure, not a full queue). Response includes `Retry-After: 5` header. Reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. As for `publish failed`, the record's id is given back so the retry can publish, and the same uncertain-publish caveat applies — the broker may already have stored the event before the timeout. | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** @@ -407,10 +407,10 @@ A `200` is returned whenever the body was read and the records were processed | 403 | `{"error":"forbidden"}` (empty-role variant: `forbidden: request has no role and no public default_role is configured`) | The resolved role lacks `insert` on the table (checked once, before any record) | | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | -| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest | +| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest — but if the failing record's publish reached the broker before failing (e.g. a client disconnect after the embedded broker had already stored it), that retry can publish a second copy of it ([#629](https://github.com/Wave-RF/WaveHouse/pull/629) closes this with an idempotency key) | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. As for `publish failed`, the failing record's id is given back and the records before it keep theirs | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | -| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time, mid-batch; includes `Retry-After: 5`. Reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. As for `publish failed`, the failing record's id is given back — but an unavailable broker that timed out may already have stored the event, so a retry can publish a second copy (the windowed-ingest follow-up closes this with an idempotency key) | +| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time, mid-batch; includes `Retry-After: 5`. Reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. As for `publish failed`, the failing record's id is given back, and the same uncertain-publish caveat applies | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] From 2c27538050e5db4ae9c7ba7a522f02eb7b313140 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:42:21 -0400 Subject: [PATCH 082/108] test(api): cover ErrUnavailable in FailedPublishReleasesTheID MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Add a table row for an mq.ErrUnavailable-wrapped publish error, asserting the 503, the released claim, and Retry-After: 5 — the sibling case to the existing ErrQueueFull row, whose own Retry-After is now asserted too. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest_test.go | 13 ++++++++----- 1 file changed, 8 insertions(+), 5 deletions(-) diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index fbff039c..17222302 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -2783,12 +2783,14 @@ func dedupHandler(t *testing.T, pub *testutil.MockPublisher, dedup dedupe.Dedupl func TestIngest_Dedup_FailedPublishReleasesTheID(t *testing.T) { t.Parallel() tests := []struct { - name string - err error - status int + name string + err error + status int + retryAfter string }{ - {"backpressure", fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull), http.StatusServiceUnavailable}, - {"other failure", errors.New("connection reset"), http.StatusInternalServerError}, + {"backpressure", fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull), http.StatusServiceUnavailable, "30"}, + {"unavailable broker", fmt.Errorf("%w: timeout", mq.ErrUnavailable), http.StatusServiceUnavailable, "5"}, + {"other failure", errors.New("connection reset"), http.StatusInternalServerError, ""}, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { @@ -2801,6 +2803,7 @@ func TestIngest_Dedup_FailedPublishReleasesTheID(t *testing.T) { w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) require.Equal(t, tt.status, w.Code) + assert.Equal(t, tt.retryAfter, w.Header().Get("Retry-After")) assert.False(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "released, not left to lapse") pub.Err = nil From cf241155afdb4ce8ea5fde8d3c39a1b796c70c7d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:42:28 -0400 Subject: [PATCH 083/108] docs(dedupe): reword Reserve's on-error guarantee for network backends MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit "On error it has released every claim it made" promises more than a network backend can keep: a timeout or a cancelled call can leave a write's outcome unknown, and that write may still land after Reserve has already returned the error. Reword to released every claim it *knows* it made, with the unknown case spelled out — the key then holds InFlight until the lease ends, like an abandoned claim. No architecture.md copy of the old wording exists to fix, and the "a failed reserve leaves nothing claimed" conformance case already tests a known (synchronous) failure, which the new wording still covers. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/dedupe.go | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/internal/dedupe/dedupe.go b/internal/dedupe/dedupe.go index ada64a5e..68061ef4 100644 --- a/internal/dedupe/dedupe.go +++ b/internal/dedupe/dedupe.go @@ -65,9 +65,11 @@ type Claim struct { // Reserve is atomic per key: of any number of concurrent Reserves for the // same key — in this process or any other sharing the backend — at most one // returns Claimed. It returns one Claim per key, in input order. On error it -// has released every claim it made (all-or-nothing from the caller's view), -// and the error wraps ErrUnavailable when retrying later can succeed -// (throttled, timed out, backend unreachable). +// has released every claim it knows it made; a write whose outcome the +// error left unknown (a timeout, a cancelled call) may still land +// afterwards, and then holds its key InFlight until the lease ends, like an +// abandoned claim. The error wraps ErrUnavailable when retrying later can +// succeed (throttled, timed out, backend unreachable). // // Commit makes Claimed claims duplicates for retention (0 = no expiry) and // ignores claims of any other status. It is unconditional: a commit that From 719a3f8e916c78bec857717408345f64cc1c5ff9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:42:34 -0400 Subject: [PATCH 084/108] perf(dedupe): skip Managed.Apply's write lock when unchanged A settings reload calls Apply for every tenant under the registry lock; on a network backend, Commit and Release can hold the read lock for as long as an outage lasts, so Apply's unconditional write lock serialized the whole reload behind them, tenant after tenant. Apply now checks under a read lock whether the desired state already holds and returns without the write lock; only a real transition takes it, re-checked once held. Covered by a test using a backend fake whose Commit blocks on a channel: a no-op Apply returns promptly while Commit is in flight, and a real transition still waits for it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/managed.go | 31 ++++++++++++++++++ internal/dedupe/managed_test.go | 57 +++++++++++++++++++++++++++++++++ 2 files changed, 88 insertions(+) diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index 0c53f97d..cd42aeb3 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -58,9 +58,24 @@ func NewManaged(open func() (Deduplicator, error)) *Managed { // already-open store stays open, an already-closed one stays closed. A // failed open leaves the store closed and returns the error — the caller // decides whether that is fatal (boot) or a logged degradation (reload). +// +// A no-op call — the desired state already holds — returns under the read +// lock alone; only a real transition takes the write lock, re-checked once +// held in case another Apply won the race. This matters because a settings +// reload calls Apply for every tenant under the registry lock: on a network +// backend, Commit and Release can hold the read lock for as long as an +// outage lasts, and the write lock waits out every reader, so an +// unconditional write lock here would serialize the whole reload behind +// them, tenant after tenant. func (m *Managed) Apply(enabled bool) error { + if m.settled(enabled) { + return nil + } m.mu.Lock() defer m.mu.Unlock() + if m.settledLocked(enabled) { + return nil + } m.enabled = enabled switch { case enabled && m.db == nil: @@ -77,6 +92,22 @@ func (m *Managed) Apply(enabled bool) error { return nil } +// settled reports whether the store already matches enabled, under its own +// read lock. +func (m *Managed) settled(enabled bool) bool { + m.mu.RLock() + defer m.mu.RUnlock() + return m.settledLocked(enabled) +} + +// settledLocked is settled's condition for a caller already holding mu (read +// or write): an already-open store while enabling, or an already-closed one +// while disabling (db is nil whenever !enabled — Apply's own invariant — so +// disabling never needs the db pointer). +func (m *Managed) settledLocked(enabled bool) bool { + return m.enabled == enabled && (!enabled || m.db != nil) +} + // Open reports whether the store is currently open. func (m *Managed) Open() bool { m.mu.RLock() diff --git a/internal/dedupe/managed_test.go b/internal/dedupe/managed_test.go index 944544e9..cbc4593b 100644 --- a/internal/dedupe/managed_test.go +++ b/internal/dedupe/managed_test.go @@ -166,3 +166,60 @@ func TestManaged_CommitAndReleaseFollowTheSwitch(t *testing.T) { require.ErrorIs(t, m.Commit(ctx, claimed, 0), ErrUnavailable) require.ErrorIs(t, m.Release(ctx, claimed), ErrUnavailable) } + +// blockingDedup's Commit blocks until unblock is closed, standing in for a +// network backend mid-outage: the caller holds Managed's read lock for as +// long as the call takes. +type blockingDedup struct { + memDedup + inCommit chan struct{} // closed once Commit is entered + unblock chan struct{} +} + +func (b *blockingDedup) Commit(ctx context.Context, claims []Claim, retention time.Duration) error { + close(b.inCommit) + <-b.unblock + return b.memDedup.Commit(ctx, claims, retention) +} + +// A no-op Apply must not queue behind an in-flight Commit: it settles under +// the read lock alone, so a reload naming the same state for every tenant +// never waits out another tenant's slow backend call. A real transition is +// the opposite — it still needs the store quiescent, so it waits for Commit +// to finish before touching it. +func TestManaged_ApplyNoOpDoesNotWaitOnCommit(t *testing.T) { + t.Parallel() + backend := &blockingDedup{memDedup: memDedup{seen: map[Key]bool{}}, inCommit: make(chan struct{}), unblock: make(chan struct{})} + m := NewManaged(func() (Deduplicator, error) { return backend, nil }) + require.NoError(t, m.Apply(true)) + + claimed := []Claim{{Key: Key{Table: "t", ID: "a"}, Status: Claimed, Token: "t"}} + commitDone := make(chan error, 1) + go func() { commitDone <- m.Commit(context.Background(), claimed, 0) }() + <-backend.inCommit // Commit is inside the backend call, holding the read lock + + noop := make(chan error, 1) + go func() { noop <- m.Apply(true) }() + select { + case err := <-noop: + require.NoError(t, err) + case <-time.After(2 * time.Second): + t.Fatal("Apply(true) blocked behind an in-flight Commit for a state that already held") + } + + // A real transition is the genuine case: it must wait for Commit, not + // race it — assert it's still pending, then let Commit finish and + // confirm Apply(false) then proceeds and closes the store. + transition := make(chan error, 1) + go func() { transition <- m.Apply(false) }() + select { + case err := <-transition: + t.Fatalf("Apply(false) returned (%v) before the in-flight Commit finished", err) + case <-time.After(50 * time.Millisecond): + } + + close(backend.unblock) + require.NoError(t, <-commitDone) + require.NoError(t, <-transition) + assert.True(t, backend.closed) +} From aea060019bf0b26badfac4acdebd6e61d41a28c8 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:42:40 -0400 Subject: [PATCH 085/108] docs(development): note dedupe's Reserve/Commit/Release in the tree The project-tree entry still described the pre-reserve CheckAndMark shape; match architecture.md's line. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/development.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 9ce612c0..b28f0b97 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -459,7 +459,7 @@ WaveHouse/ │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration │ ├── coord/ # Leases with fencing tokens (in-process Local, RunElected, coordtest suite) -│ ├── dedupe/ # Optional deduplication (Pebble) +│ ├── dedupe/ # Optional deduplication (Reserve/Commit/Release; Pebble) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper │ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens, dedupe keys) From f177f4d011e135da427c943546ffd92f47598a56 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:54:18 -0400 Subject: [PATCH 086/108] fix(test): remove the broker store after late consumer writes land The embedded NATS server writes each durable consumer's state (obs//o.dat, via a temp file renamed into place) from a goroutine neither Shutdown nor WaitForShutdown joins; its consumer store waits for it at close only while state is unwritten, and for at most 100ms. A write already under way therefore lands after EmbeddedNATS.Close returns, and t.TempDir's one-shot RemoveAll met the late entry as "directory not empty". The worker's own teardown was already joined; the late writer is the server's. storedir.New(t) is now the store directory for every test that puts a broker on disk (testutil.NewEmbeddedMQ, the mq and mqtest suites, the app tests' data_dir, three integration tests). Its cleanup runs after the broker's Close and removes the store again whenever a directory was refilled between being read and being removed. Each late write adds at most two entries and none once its directory is gone, so the removal ends without a timer. It replaces two sleep-and-retry copies in the mq tests. A rename-aside fence was tried first and does not work: the flusher's rename resolves its paths before the fence and lands after it. TestStartIngestWorker_StopFunc_RespectsShutdownDeadline now joins the worker its deadline abandons before the broker closes. Under six parallel race-enabled test processes on 14 CPU hogs, the dispatch-loop and StartIngestWorker tests failed 78 of 1800 runs before and 0 of 1800 after. With the helper instrumented to log its retries, 102 of 1800 teardowns needed a second removal and every one succeeded on it. Closes #442 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- AGENTS.md | 2 +- CHANGELOG.md | 1 + docs/src/content/docs/development.md | 2 +- internal/app/app_test.go | 3 +- internal/app/roles_test.go | 3 +- internal/ingest/worker_test.go | 11 +++-- internal/mq/embedded.go | 3 +- internal/mq/embedded_test.go | 49 +++++++-------------- internal/mq/mqtest/embedded_test.go | 25 +---------- internal/testutil/storedir/storedir.go | 47 ++++++++++++++++++++ internal/testutil/storedir/storedir_test.go | 28 ++++++++++++ internal/testutil/testutil.go | 7 +-- tests/integration/ingest_outage_test.go | 3 +- tests/integration/query_errors_test.go | 3 +- tests/integration/tenants_test.go | 3 +- 15 files changed, 118 insertions(+), 72 deletions(-) create mode 100644 internal/testutil/storedir/storedir.go create mode 100644 internal/testutil/storedir/storedir_test.go diff --git a/AGENTS.md b/AGENTS.md index f6580a22..dcc13cc9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -446,7 +446,7 @@ internal/query/ → Structured query AST + SQL builder internal/settings/ → Settings directory (validate, adopted snapshot + reload, watcher, embedded seed) internal/stream/ → SSE fan-out (event Hub: project once per role, Subscriber outbound queue, Bucket fan-out, keepalive Heartbeater wheel) internal/tenant/ → Tenant id (type, grammar, reserved default, request header name) -internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger) +internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; storedir/ is the embedded broker's store directory in tests, removed once late consumer-state writes land) tests/ → Integration & E2E tests tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer) tests/e2e/ → E2E test stack (scripts/orchestrator boots a ClickHouse testcontainer + the wavehouse-cov binary) diff --git a/CHANGELOG.md b/CHANGELOG.md index ccdd58bb..a4f91f32 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,6 +85,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed +- **Tests that start the embedded broker no longer fail removing its store after passing** (`internal/testutil/storedir` (new, + tests), `internal/testutil/testutil.go`, `internal/mq/embedded.go` (comment), `internal/mq/{embedded_test,mqtest/embedded_test}.go`, `internal/ingest/worker_test.go`, `internal/app/{app,roles}_test.go`, `tests/integration/{ingest_outage,query_errors,tenants}_test.go`, `AGENTS.md`, `docs/src/content/docs/development.md`): [#442](https://github.com/Wave-RF/WaveHouse/issues/442). The NATS server writes each durable consumer's state (`obs//o.dat`, through a temporary file renamed into place) from a goroutine that neither `Shutdown` nor `WaitForShutdown` joins, and its consumer store waits for that goroutine at close only when state is still unwritten, for at most 100ms — so a write already under way lands after `EmbeddedNATS.Close` returns, and `t.TempDir`'s one-shot `RemoveAll` met the late entry as `directory not empty`. Under parallel test processes it failed about 4% of the ingest worker tests (78 of 1,800 runs). Every store a test puts on disk now comes from `storedir.New(t)`, whose cleanup — after the broker's `Close` — removes it again whenever a directory was refilled between being read and being removed: each late write adds at most two entries and none once its directory is gone, so the removal ends without a timer (0 of 1,800 under the same load). It replaces two sleep-and-retry copies in the `internal/mq` tests. `TestStartIngestWorker_StopFunc_RespectsShutdownDeadline` also joins the worker its deadline abandons before the broker closes, rather than leaving it to ack on a closed connection. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 6ee7ceed..d16835f5 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -347,7 +347,7 @@ Each test target writes `covdata` to `tmp/coverage//data/`, renders a tex - **Unit tests** live beside the code they test (e.g., `internal/discovery/discovery_test.go`). They use mocks or embedded NATS (in-process, no Docker needed). - **Integration tests** use the `//go:build integration` build tag. `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. -Shared test utilities live in `internal/testutil/`. The packages log through `slog.Default()`, so tests reach log output through `internal/testutil/logtest`: `logtest.Silence()` in a package's `TestMain` discards it, and `logtest.Capture(t, level)` routes it to a buffer for a test that asserts on log lines — such a test must not call `t.Parallel()`, because the default logger is process-wide. +Shared test utilities live in `internal/testutil/`. The packages log through `slog.Default()`, so tests reach log output through `internal/testutil/logtest`: `logtest.Silence()` in a package's `TestMain` discards it, and `logtest.Capture(t, level)` routes it to a buffer for a test that asserts on log lines — such a test must not call `t.Parallel()`, because the default logger is process-wide. A test that starts the embedded broker keeps its store in `internal/testutil/storedir`'s `storedir.New(t)` rather than a bare `t.TempDir()` (`testutil.NewEmbeddedMQ` does): the NATS server can finish writing a consumer's state after `Close` returns, which fails `t.TempDir`'s one-shot removal, and `storedir` removes the store again until those writes have landed ([#442](https://github.com/Wave-RF/WaveHouse/issues/442)). ### Adding New Tests diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 6a8bc449..7af8c8f3 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -35,6 +35,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil" "github.com/Wave-RF/WaveHouse/internal/testutil/logtest" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // None of these tests run in parallel: New installs a process-wide default @@ -94,7 +95,7 @@ func writeSettings(t *testing.T, patch map[string]any) string { func testConfig(t *testing.T, settingsDir string) *config.Config { t.Helper() return &config.Config{ - DataDir: t.TempDir(), + DataDir: storedir.New(t), Server: config.Server{Port: closedPort(t), ShutdownTimeout: 2}, MQ: config.MQ{Backend: config.MQEmbedded}, Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, diff --git a/internal/app/roles_test.go b/internal/app/roles_test.go index 48fb0fa1..c50eb9a3 100644 --- a/internal/app/roles_test.go +++ b/internal/app/roles_test.go @@ -17,6 +17,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/config" "github.com/Wave-RF/WaveHouse/internal/coord" "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // Each role wires its own components and nothing else; the settings registry, @@ -109,7 +110,7 @@ func TestNew_OpsOnlyRouter(t *testing.T) { sweeperCfg := *cfg sweeperCfg.Roles = []config.Role{config.RoleSweeper} // Its own store: full's embedded JetStream is still open on cfg.DataDir. - sweeperCfg.DataDir = t.TempDir() + sweeperCfg.DataDir = storedir.New(t) a := newApp(t, &sweeperCfg, Options{}) for _, path := range []string{"/livez", "/readyz", "/healthz", "/version"} { diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index f07500bc..7b884093 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -254,10 +254,7 @@ func TestStartIngestWorker_StopFunc_RespectsShutdownDeadline(t *testing.T) { <-release w.WriteHeader(http.StatusOK) })) - t.Cleanup(func() { - close(release) - chSrv.Close() - }) + t.Cleanup(chSrv.Close) u, _ := url.Parse(chSrv.URL) host, port, _ := net.SplitHostPort(u.Host) @@ -269,6 +266,12 @@ func TestStartIngestWorker_StopFunc_RespectsShutdownDeadline(t *testing.T) { return chconn.Target{URL: fmt.Sprintf("http://%s:%s", host, port), Username: "u", Password: "p", Database: "db"} }, nil) require.NoError(t, err) + // The deadline below abandons the worker, not its insert: finish the insert + // and join the worker before the broker closes under its ack. + t.Cleanup(func() { + close(release) + assert.NoError(t, stopFn(context.Background())) + }) // Publish so there's an in-flight insert blocking on `release`. err = emb.Publish(ctx, mq.Topic{Tenant: tenant.Default, Table: "events"}, makeEnvelope(t, "events", "", map[string]any{"id": 1})) diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 15ca485f..c8c7581d 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -1086,7 +1086,8 @@ func (e *EmbeddedNATS) Close() error { // Owning the lifecycle (NoSigs, #287) means waiting it out: without this, // run()'s remaining defers unwind while JetStream is still tearing down // and the process can exit mid-shutdown (as-if-crashed stream state). - // Milliseconds for an in-process server. + // Milliseconds for an in-process server. It does not join a durable's + // state flusher, whose write under way can land after Close returns (#442). e.server.WaitForShutdown() return nil } diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 55e7fec4..e172c903 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -10,6 +10,7 @@ import ( "time" "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" "github.com/nats-io/nats.go" "github.com/nats-io/nats.go/jetstream" "github.com/stretchr/testify/assert" @@ -19,26 +20,6 @@ import ( // testBudget is the byte budget newTestEmbedded opens each queue at. const testBudget = 64 << 20 -// storeDir is a temporary directory for a broker's store whose removal -// retries briefly: a consumer's state file can land after Close has returned, -// which fails t.TempDir's one-shot RemoveAll (#442). The retrying cleanup runs -// first (cleanups are LIFO), leaving t.TempDir an empty directory to remove. -func storeDir(t *testing.T) string { - t.Helper() - dir := filepath.Join(t.TempDir(), "store") - t.Cleanup(func() { - var err error - for range 50 { - if err = os.RemoveAll(dir); err == nil { - return - } - time.Sleep(20 * time.Millisecond) - } - t.Errorf("remove %s: %v", dir, err) - }) - return dir -} - // openEmbedded starts an EmbeddedNATS over dir, closed by the test framework. func openEmbedded(t *testing.T, dir string) *EmbeddedNATS { t.Helper() @@ -53,7 +34,7 @@ func openEmbedded(t *testing.T, dir string) *EmbeddedNATS { // at testBudget. func newTestEmbedded(t *testing.T, tenants ...tenant.ID) *EmbeddedNATS { t.Helper() - e := openEmbedded(t, storeDir(t)) + e := openEmbedded(t, storedir.New(t)) if len(tenants) == 0 { tenants = []tenant.ID{tenant.Default} } @@ -167,7 +148,7 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // at a tenth of it, dropping its oldest when full. No other tenant gets one. func TestEmbeddedNATS_SetMaxBytes_OpensTheTenantsQueue(t *testing.T) { t.Parallel() - e := openEmbedded(t, storeDir(t)) + e := openEmbedded(t, storedir.New(t)) assert.Zero(t, e.MaxBytes("acme"), "no budget applied yet") require.NoError(t, e.SetMaxBytes(t.Context(), "acme", testBudget)) @@ -342,7 +323,7 @@ func TestEmbeddedNATS_DefaultLogger(t *testing.T) { t.Parallel() // NewEmbedded without a logger should not panic — it falls back to the // default slog logger. - e, err := NewEmbedded(storeDir(t)) + e, err := NewEmbedded(storedir.New(t)) require.NoError(t, err) t.Cleanup(func() { _ = e.Close() }) } @@ -458,7 +439,7 @@ func TestEmbeddedNATS_SetMaxBytes_IngestFailureChangesNothing(t *testing.T) { // however recently a publish tried. func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { t.Parallel() - dir := storeDir(t) + dir := storedir.New(t) // The dead-letter stream is the first of the pair to open. A failed open // removes what was in the way, so the obstacle is put back before each // attempt meant to fail. @@ -509,7 +490,7 @@ func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { // resize and reload takes. Once the window has passed, a publish tries again. func TestEmbeddedNATS_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T) { t.Parallel() - dir := storeDir(t) + dir := storedir.New(t) block := filepath.Join(dir, "jetstream", "$G", "streams", dlqStreamName("acme")) obstruct := func() { t.Helper() @@ -574,7 +555,7 @@ func TestEmbeddedNATS_PacesTheRetriesOfAQueueThatCannotOpen(t *testing.T) { // joined, so its row reaches them rather than a stream nobody reads. func TestEmbeddedNATS_Publish_OpensAQueueItsOpenGaveUpOn(t *testing.T) { t.Parallel() - dir := storeDir(t) + dir := storedir.New(t) block := filepath.Join(dir, "jetstream", "$G", "streams", ingestStreamName("acme")) require.NoError(t, os.MkdirAll(filepath.Dir(block), 0o750)) require.NoError(t, os.WriteFile(block, nil, 0o600)) @@ -618,7 +599,7 @@ func TestEmbeddedNATS_SetMaxBytes_UndoRestoresTheIngestStreamsCap(t *testing.T) t.Parallel() ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - dir := storeDir(t) + dir := storedir.New(t) first, err := NewEmbedded(dir) require.NoError(t, err) require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) @@ -641,7 +622,7 @@ func TestEmbeddedNATS_SetMaxBytes_UndoRestoresTheIngestStreamsCap(t *testing.T) // queue itself is open, so SetMaxBytes succeeds. func TestEmbeddedNATS_Consume_ReportsAQueueItCannotJoin(t *testing.T) { t.Parallel() - e := openEmbedded(t, storeDir(t)) + e := openEmbedded(t, storedir.New(t)) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() // A durable name the client refuses: with no queue yet, nothing checks it. @@ -666,7 +647,7 @@ func TestEmbeddedNATS_Consume_ReportsAQueueItCannotJoin(t *testing.T) { // stream keeps what it holds, capped at that, and every row survives. func TestEmbeddedNATS_SetMaxBytes_NeverShrinksTheDeadLetterQueueBelowWhatItHolds(t *testing.T) { t.Parallel() - e := openEmbedded(t, storeDir(t)) + e := openEmbedded(t, storedir.New(t)) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() require.NoError(t, e.SetMaxBytes(ctx, "acme", 10<<20)) @@ -881,7 +862,7 @@ func TestEmbeddedNATS_DeadLetter_ReopensAMissingQueue(t *testing.T) { func TestEmbeddedNATS_Publish_QueueFull(t *testing.T) { t.Parallel() - e := openEmbedded(t, storeDir(t)) + e := openEmbedded(t, storedir.New(t)) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() require.NoError(t, e.SetMaxBytes(ctx, "acme", 4<<10)) @@ -1364,7 +1345,7 @@ func TestNewEmbedded_AStoreItCannotCreateFailsAtOnce(t *testing.T) { // so no tenant's queue could open beside them. func TestNewEmbedded_DeletesTheStreamsAnEarlierBuildShared(t *testing.T) { t.Parallel() - dir := storeDir(t) + dir := storedir.New(t) old, err := NewEmbedded(dir) require.NoError(t, err) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) @@ -1395,7 +1376,7 @@ func TestNewEmbedded_ASplitPairIsAppliedAgainAtBoot(t *testing.T) { t.Parallel() ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - dir := storeDir(t) + dir := storedir.New(t) first, err := NewEmbedded(dir) require.NoError(t, err) for _, id := range []tenant.ID{"split", "gone", "guarded"} { @@ -1433,7 +1414,7 @@ func TestNewEmbedded_ASplitPairIsAppliedAgainAtBoot(t *testing.T) { // so what such a tenant had queued still reaches the worker. func TestNewEmbedded_TakesStockOfTheQueuesOnDisk(t *testing.T) { t.Parallel() - dir := storeDir(t) + dir := storedir.New(t) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() first, err := NewEmbedded(dir) @@ -1484,7 +1465,7 @@ func TestEmbeddedNATS_ADurableOnDiskIsReusedAcrossARestart(t *testing.T) { } { t.Run(tt.name, func(t *testing.T) { t.Parallel() - dir := storeDir(t) + dir := storedir.New(t) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() topic := Topic{Tenant: "acme", Table: "t"} diff --git a/internal/mq/mqtest/embedded_test.go b/internal/mq/mqtest/embedded_test.go index 98b619d9..f3be090a 100644 --- a/internal/mq/mqtest/embedded_test.go +++ b/internal/mq/mqtest/embedded_test.go @@ -3,21 +3,19 @@ package mqtest_test import ( - "os" - "path/filepath" "testing" - "time" "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/mq/mqtest" "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" "github.com/stretchr/testify/require" ) func TestEmbeddedNATS_Conformance(t *testing.T) { mqtest.Run(t, mqtest.Harness{ New: func(t *testing.T) mq.Broker { - e, err := mq.NewEmbedded(storeDir(t)) + e, err := mq.NewEmbedded(storedir.New(t)) require.NoError(t, err) t.Cleanup(func() { _ = e.Close() }) for _, id := range []tenant.ID{mqtest.Acme, mqtest.Globex} { @@ -55,22 +53,3 @@ func TestEmbeddedNATS_Conformance(t *testing.T) { }, }) } - -// storeDir is a temporary store directory whose removal retries briefly: under -// parallel load a consumer's state file can land after Close has returned, -// which fails t.TempDir's one-shot RemoveAll. The retrying cleanup runs first -// (cleanups are LIFO), leaving t.TempDir an empty directory to remove. -func storeDir(t *testing.T) string { - dir := filepath.Join(t.TempDir(), "store") - var err error - t.Cleanup(func() { - for range 50 { - if err = os.RemoveAll(dir); err == nil { - return - } - time.Sleep(20 * time.Millisecond) - } - t.Errorf("remove %s: %v", dir, err) - }) - return dir -} diff --git a/internal/testutil/storedir/storedir.go b/internal/testutil/storedir/storedir.go new file mode 100644 index 00000000..aae61863 --- /dev/null +++ b/internal/testutil/storedir/storedir.go @@ -0,0 +1,47 @@ +// Package storedir gives a test a directory for the embedded message broker's +// store. It imports nothing from the repository, so internal/mq's own tests can +// use it. +package storedir + +import ( + "errors" + "os" + "path/filepath" + "syscall" + "testing" +) + +// maxRemovals bounds the store's removal: a consumer-state write still under way +// when the broker closes adds at most two entries after it — its temporary file, +// then the rename into place — so a removal is refilled at most twice per +// durable. Past this many, something is writing that Close did not stop. +const maxRemovals = 32 + +// New returns an empty directory under t.TempDir for a broker's store, removed +// once the broker is closed: New's cleanup runs after the test's own, Close +// among them (cleanups run last-in, first-out). +// +// A closed broker's store is not yet quiescent: the embedded NATS server writes +// each durable consumer's state from a goroutine its Shutdown does not join, so +// a write under way can land after Close has returned — failing t.TempDir's +// one-shot RemoveAll with "directory not empty" (#442). There is nothing to +// wait on, so the removal is tried again whenever a directory was refilled +// between reading and removing it. That ends without a clock: those writes add a +// bounded number of entries, and none once their directory is gone. +func New(t testing.TB) string { + t.Helper() + dir := filepath.Join(t.TempDir(), "store") + if err := os.Mkdir(dir, 0o700); err != nil { + t.Fatalf("store directory: %v", err) + } + t.Cleanup(func() { + err := os.RemoveAll(dir) + for i := 1; i < maxRemovals && errors.Is(err, syscall.ENOTEMPTY); i++ { + err = os.RemoveAll(dir) + } + if err != nil { + t.Errorf("remove the store: %v", err) + } + }) + return dir +} diff --git a/internal/testutil/storedir/storedir_test.go b/internal/testutil/storedir/storedir_test.go new file mode 100644 index 00000000..6720a534 --- /dev/null +++ b/internal/testutil/storedir/storedir_test.go @@ -0,0 +1,28 @@ +package storedir + +import ( + "os" + "path/filepath" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +func TestNew_RemovedAfterTheTestsOwnCleanups(t *testing.T) { + var dir string + t.Run("store", func(t *testing.T) { + dir = New(t) + entries, err := os.ReadDir(dir) + require.NoError(t, err) + assert.Empty(t, entries, "the store starts empty") + // Registered after New, as a broker's Close is, so it runs first and + // what it writes is removed with the rest. + t.Cleanup(func() { + obs := filepath.Join(dir, "jetstream", "obs") + require.NoError(t, os.MkdirAll(obs, 0o700)) + require.NoError(t, os.WriteFile(filepath.Join(obs, "o.dat"), nil, 0o600)) + }) + }) + assert.NoDirExists(t, dir) +} diff --git a/internal/testutil/testutil.go b/internal/testutil/testutil.go index db119686..8b3cefb4 100644 --- a/internal/testutil/testutil.go +++ b/internal/testutil/testutil.go @@ -16,6 +16,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/discovery" "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // NewTestSchemaRegistry creates a SchemaRegistry pre-loaded with the given @@ -40,13 +41,13 @@ func NewTestSchemaRegistry(t testing.TB, tables []*discovery.TableSchema) *disco // hardcoding the same literal twice. const TestServerVersion = "24.8.1.1" -// NewEmbeddedMQ starts the embedded broker over a temporary directory, closed -// by the test framework, with a queue open for each of tenants — +// NewEmbeddedMQ starts the embedded broker over a storedir.New directory, +// closed by the test framework, with a queue open for each of tenants — // tenant.Default when none is named — at maxBytes: a tenant has a queue once // its budget is applied, as the wiring does for every tenant it serves. func NewEmbeddedMQ(t testing.TB, maxBytes int64, tenants ...tenant.ID) *mq.EmbeddedNATS { t.Helper() - emb, err := mq.NewEmbedded(t.TempDir()) + emb, err := mq.NewEmbedded(storedir.New(t)) require.NoError(t, err) t.Cleanup(func() { _ = emb.Close() }) if len(tenants) == 0 { diff --git a/tests/integration/ingest_outage_test.go b/tests/integration/ingest_outage_test.go index 752d84b2..53faa2d0 100644 --- a/tests/integration/ingest_outage_test.go +++ b/tests/integration/ingest_outage_test.go @@ -18,6 +18,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // TestIngest_ClickHouseOutage_RetriedNotDeadLettered stops a real ClickHouse @@ -39,7 +40,7 @@ func TestIngest_ClickHouseOutage_RetriedNotDeadLettered(t *testing.T) { const table = "outage_events" require.NoError(t, ch.conn.Exec(ctx, "CREATE TABLE "+table+" (id UInt32) ENGINE = MergeTree ORDER BY id")) - broker, err := mq.NewEmbedded(t.TempDir()) + broker, err := mq.NewEmbedded(storedir.New(t)) require.NoError(t, err) t.Cleanup(func() { _ = broker.Close() }) require.NoError(t, broker.SetMaxBytes(ctx, tenant.Default, 64<<20)) diff --git a/tests/integration/query_errors_test.go b/tests/integration/query_errors_test.go index 9eccd851..d29660ef 100644 --- a/tests/integration/query_errors_test.go +++ b/tests/integration/query_errors_test.go @@ -22,6 +22,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/app" "github.com/Wave-RF/WaveHouse/internal/chconn" "github.com/Wave-RF/WaveHouse/internal/config" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // queryError is the error envelope a failed ClickHouse query answers with. @@ -115,7 +116,7 @@ func TestQueryErrors_ClickHouseDown(t *testing.T) { ln, err := lc.Listen(ctx, "tcp", "127.0.0.1:0") require.NoError(t, err) cfg := &config.Config{ - DataDir: t.TempDir(), + DataDir: storedir.New(t), Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, MQ: config.MQ{Backend: config.MQEmbedded}, diff --git a/tests/integration/tenants_test.go b/tests/integration/tenants_test.go index 16d888ba..e3c756b6 100644 --- a/tests/integration/tenants_test.go +++ b/tests/integration/tenants_test.go @@ -20,6 +20,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/app" "github.com/Wave-RF/WaveHouse/internal/config" "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // TestNestedDirectory_PerTenantPoolsAndDiscovery boots the real wiring over @@ -58,7 +59,7 @@ func TestNestedDirectory_PerTenantPoolsAndDiscovery(t *testing.T) { ln, err := lc.Listen(ctx, "tcp", "127.0.0.1:0") require.NoError(t, err) cfg := &config.Config{ - DataDir: t.TempDir(), + DataDir: storedir.New(t), Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, Auth: config.Auth{OperatorKey: operatorKey}, From 5c45de2b38622c1bc7ec0c0a116532e275c34281 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:18:51 -0400 Subject: [PATCH 087/108] docs(agents): name Reserve/Commit/Release in the dedupe tree line Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- AGENTS.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/AGENTS.md b/AGENTS.md index 8e235315..73eb427f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -434,7 +434,7 @@ internal/chconn/ → ClickHouse pools, one per connection tuple among the internal/chsql/ → Shared ClickHouse SQL helpers (identifier quoting + bind-safety) internal/config/ → Configuration structs + loader internal/coord/ → Leases with fencing tokens (interface, in-process Local, RunElected, coordtest conformance suite) -internal/dedupe/ → Optional deduplication (interface + embedded/distributed) +internal/dedupe/ → Optional deduplication (Reserve/Commit/Release interface + embedded Pebble) internal/discovery/ → ClickHouse schema introspection + ingest validation internal/ingest/ → Batch buffer with DLQ + Active Sweeper (NATS message lifecycle) internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache namespace tokens, dedupe keys) From 3513823459cec691fb42d7105d39d94bf545b447 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:27:10 -0400 Subject: [PATCH 088/108] fix(config): cap the embedded-mq dedupe lease by lease+ceil(lease)+1s The old check allowed a lease up to 59.5s (2*lease+1s <= 2m), but the DynamoDB claim it bounds rounds its expiry up to the next whole second, and the in-flight 503 it drives can go out a second late. The true worst case is lease + ceil(lease) + 1s, which only stays under the embedded queue's 2-minute duplicate window through 59s exactly: 59.5s (and anything else over 59s) already crosses it once ceil(lease) steps to the next second. Co-Authored-By: Claude Opus 5.5 (1M context) --- internal/config/backends.go | 31 ++++++++++++++++++++++++------- internal/config/backends_test.go | 10 +++++----- 2 files changed, 29 insertions(+), 12 deletions(-) diff --git a/internal/config/backends.go b/internal/config/backends.go index 42d7b705..b2dc0dd0 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -170,11 +170,22 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { // counted from the stored publish. const embeddedDuplicateWindow = 2 * time.Minute -// maxEmbeddedLease is the longest dedupe.lease that window covers: a client -// that obeys the in-flight 503's Retry-After (the whole lease) after a -// publish whose outcome it never learned republishes up to twice the lease -// after the claim, and DynamoDB rounds a claim's expiry up to the second. -const maxEmbeddedLease = (embeddedDuplicateWindow - time.Second) / 2 +// maxEmbeddedLease is the longest dedupe.lease the duplicate window covers — +// the largest whole second satisfying the rule below. It is informational +// only: validateBackends checks the rule itself, not this constant, since +// the rule's ceiling steps at each whole second rather than moving linearly +// with the lease. +const maxEmbeddedLease = 59 * time.Second + +// ceilSecond rounds d up to the next whole second, as a DynamoDB claim's +// expiry does (epoch seconds, rounded up) — so a claim taken out just before +// the tick it is stamped with can stay live up to a second past the lease. +func ceilSecond(d time.Duration) time.Duration { + if r := d % time.Second; r != 0 { + d += time.Second - r + } + return d +} // validateBackends checks every layer's backend and its sub-block, then the // rules that span two layers. @@ -184,8 +195,14 @@ func (c *Config) validateBackends() error { return err } } - if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > maxEmbeddedLease { - return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s is over %s with the embedded mq: twice the lease plus 1s must fit its %s duplicate window, since a client obeying the in-flight 503's Retry-After republishes up to twice the lease after the claim", c.Dedupe.Lease, maxEmbeddedLease, embeddedDuplicateWindow) + // A client obeying the in-flight 503's Retry-After (the whole lease) + // republishes at t0+lease at the earliest. But a claim can outlive its + // own lease by up to a second (DynamoDB rounds expiry up to the second), + // so the last such 503 can go out at t0+lease+1s, and the republish it + // asks for lands at t0+lease+1s+ceil(lease). That must still fall inside + // the embedded duplicate window: lease + ceil(lease) + 1s <= 2m. + if worst := c.Dedupe.Lease + ceilSecond(c.Dedupe.Lease) + time.Second; c.MQ.Backend == MQEmbedded && worst > embeddedDuplicateWindow { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s is over %s with the embedded mq: lease + ceil(lease) + 1s (%s) must fit its %s duplicate window, since a client obeying the in-flight 503's Retry-After can republish that late", c.Dedupe.Lease, maxEmbeddedLease, worst, embeddedDuplicateWindow) } return nil } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index f4c01b0c..fd2005a8 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -259,8 +259,8 @@ func TestValidate_Dedupe(t *testing.T) { c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true }, ""}, {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB = DedupeDynamoDBConfig{CreateTable: true} }, ""}, - {"lease just under a minute", func(c *Config) { c.Dedupe.Lease = 59 * time.Second }, ""}, - {"lease at the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond }, ""}, + {"lease at the cap", func(c *Config) { c.Dedupe.Lease = 59 * time.Second }, ""}, + {"lease just past the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 100*time.Millisecond }, "is over 59s with the embedded mq"}, {"create_table without an endpoint", func(c *Config) { dynamo(c) c.Dedupe.DynamoDB.CreateTable = true @@ -276,9 +276,9 @@ func TestValidate_Dedupe(t *testing.T) { {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0"}, {"zero concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = 0 }, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0"}, {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, - {"lease of a minute", func(c *Config) { c.Dedupe.Lease = time.Minute }, "dedupe.lease (WH_DEDUPE_LEASE) 1m0s is over 59.5s with the embedded mq: twice the lease plus 1s must fit its 2m0s duplicate window"}, - {"lease just past the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond + 1 }, "is over 59.5s with the embedded mq"}, - {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, "is over 59.5s with the embedded mq"}, + {"lease of a minute", func(c *Config) { c.Dedupe.Lease = time.Minute }, "dedupe.lease (WH_DEDUPE_LEASE) 1m0s is over 59s with the embedded mq: lease + ceil(lease) + 1s (2m1s) must fit its 2m0s duplicate window"}, + {"lease at the old 59.5s cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond }, "is over 59s with the embedded mq"}, + {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, "is over 59s with the embedded mq"}, } for _, tc := range cases { t.Run(tc.name, func(t *testing.T) { From 2c42f0add33c1b5e9ebfbb1d929b5f9298db91d3 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:27:31 -0400 Subject: [PATCH 089/108] docs(dedupe): sweep the lease cap, the table check's deadline, and what a reload waits on MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every remaining 59.5s (config.yaml, CHANGELOG.md, configuration.mdx, architecture.md) now reads 59s with the lease+ceil(lease)+1s rule spelled out, matching the previous commit's config change. Adds the boot/background table check's own 10x-timeout deadline (2.5s by default) to the dedupe.dynamodb.timeout row, since that call is not one of the per-request calls the row otherwise describes. Marks the development.md tree line "Pebble or DynamoDB" now that this PR makes the backend selectable, replacing the prior "not yet selectable at boot" framing that architecture.md's dynamodb.go entry already dropped. Qualifies "a reload never waits on the table" everywhere it appears (configuration.mdx, deployment.md, CHANGELOG.md, architecture.md's wire.go paragraph): true for a tenant whose dedupe.enabled did not change, thanks to Managed.Apply's no-op fast path settling under a read lock alone — but switching a tenant's dedupe off is a genuine transition, which takes the write lock and so waits for that tenant's in-flight Reserve/Commit/Release calls to finish first. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 4 ++-- docs/src/content/docs/configuration.mdx | 8 ++++---- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/development.md | 2 +- 6 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5ec73acb..747fbd8f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload never waits on the table: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; switching a tenant's dedupe off is the exception, waiting for that tenant's in-flight calls to finish before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/config.yaml b/config.yaml index 3d24dc79..b45d5b9d 100644 --- a/config.yaml +++ b/config.yaml @@ -57,7 +57,7 @@ mq: backend: embedded # NATS JetStream under /nats dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) - lease: 30s # how long a claimed id stays pending; at most 59.5s with the embedded mq (2*lease + 1s within its 2m duplicate window) + lease: 30s # how long a claimed id stays pending; at most 59s with the embedded mq (lease + ceil(lease) + 1s within its 2m duplicate window) reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend, and DynamoDB's idle connections per host; the fan-out has no effect yet (ingest sends one id per call) # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 83b54a81..6f18bb5a 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -93,7 +93,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads: it applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -120,7 +120,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: while `mq.backend` is `embedded`, twice `dedupe.lease` plus one second must fit the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`), a cap of 59.5s (`maxEmbeddedLease`), because a client obeying the in-flight `503`'s `Retry-After` republishes up to twice the lease after the claim and DynamoDB rounds a claim's expiry up to the second. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: while `mq.backend` is `embedded`, `dedupe.lease` plus its own ceiling to the next whole second (`ceilSecond`) plus one more second must fit the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`), a cap of 59s (`maxEmbeddedLease`), because a client obeying the in-flight `503`'s `Retry-After` can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index d89d1833..a1286b53 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -56,19 +56,19 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, twice the lease plus one second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59.5s`: a client that obeys `Retry-After` after a publish whose outcome it never learned republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, the lease plus its own ceiling to the next whole second plus one more second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59s`: a client that obeys `Retry-After` after a publish whose outcome it never learned can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | | `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match. Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload never waits on the table: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. Switching a tenant's dedupe off is the exception — it waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish before the store closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(required)* | The shared table. | | `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`); no region from either refuses boot. | | `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | -| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. The retries back off with full jitter, each wait capped at `timeout / (2 × (max_attempts − 1))`, so together they wait at most half of it and a throttled call fails on its last attempt's answer rather than on the deadline. `0` refuses boot. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. The retries back off with full jitter, each wait capped at `timeout / (2 × (max_attempts − 1))`, so together they wait at most half of it and a throttled call fails on its last attempt's answer rather than on the deadline. The boot and background table check (verifying the key schema and TTL) is not one of these calls: it runs under its own deadline of 10 × `timeout` (`2.5s` by default). `0` refuses boot. | | `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. More attempts share the same half of `timeout` for their waits, so each retry waits less rather than the call running longer. `0` refuses boot. | | `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. Anything else, empty included, refuses boot. | | `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | @@ -265,7 +265,7 @@ cache: dedupe: backend: pebble # in-process Pebble under /pebble; or dynamodb - lease: 30s # at most 59.5s with the embedded mq + lease: 30s # at most 59s with the embedded mq reserve_concurrency: 64 # dynamodb: # read only when backend is dynamodb # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index de5fb64b..e9aca371 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload, which never waits on the table). No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; switching a tenant's dedupe off waits for that tenant's in-flight calls to finish before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index ab97cb79..4ca7a43d 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -459,7 +459,7 @@ WaveHouse/ │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration │ ├── coord/ # Leases with fencing tokens (in-process Local, RunElected, coordtest suite) -│ ├── dedupe/ # Optional deduplication (Reserve/Commit/Release; Pebble, DynamoDB) +│ ├── dedupe/ # Optional deduplication (Reserve/Commit/Release; Pebble or DynamoDB) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper │ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens, dedupe keys) From d0a61e4a93554edc0a8e625f22103f90c1ec173a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:27:45 -0400 Subject: [PATCH 090/108] test(app): a reload must not wait behind a tenant's in-flight DynamoDB call MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds a fake per-op hang (fakeDynamo.setHangOn) alongside the existing whole-endpoint one, then exercises Managed.Apply's no-op fast path end to end: a tenant with dedupe on has a Commit in flight against a BatchWriteItem that never answers, and asserts that a reload naming no change for that tenant still returns in well under a second, and that a concurrent Reserve for the same tenant is not blocked either — which it would be if the reload's Apply took the write lock unconditionally, since a pending writer blocks new readers too. Verified by hand: temporarily skipping the settled check in internal/dedupe/managed.go's Apply made the reload assertion fail at 2.475s (Commit's own retry/backoff giving up on the hung call, not a deadlock) instead of passing in 0.02–0.03s; reverted, `git diff` on that file confirmed clean before this commit. Co-Authored-By: Claude Opus 5.5 (1M context) --- internal/app/dedupe_dynamodb_test.go | 64 ++++++++++++++++++++++++++-- 1 file changed, 61 insertions(+), 3 deletions(-) diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index f3fb1e59..71bf294c 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -23,11 +23,13 @@ import ( // fakeDynamo answers the DynamoDB JSON protocol for one table, enough for // boot's check, the dev create path, and a claim and its commit. Whether the -// table exists, and whether the endpoint hangs, are the test's to switch. +// table exists, and whether the endpoint hangs (every call, or one op +// alone), are the test's to switch. type fakeDynamo struct { mu sync.Mutex exists bool hangs bool + hangOn string // hang calls of this op alone, once set; "" hangs none this way calls []string } @@ -43,6 +45,12 @@ func (f *fakeDynamo) setHangs(v bool) { f.hangs = v } +func (f *fakeDynamo) setHangOn(op string) { + f.mu.Lock() + defer f.mu.Unlock() + f.hangOn = op +} + func (f *fakeDynamo) called(op string) bool { return f.count(op) > 0 } func (f *fakeDynamo) count(op string) int { @@ -58,6 +66,9 @@ func (f *fakeDynamo) count(op string) int { } func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { + // Drained before any hang below: with the body unread, an SDK write + // deadline or the client giving up never reaches this handler, since the + // connection looks like it's still waiting for us to consume it. _, _ = io.Copy(io.Discard, r.Body) _, op, _ := strings.Cut(r.Header.Get("X-Amz-Target"), ".") f.mu.Lock() @@ -65,9 +76,9 @@ func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { if op == "CreateTable" { f.exists = true } - exists, hangs := f.exists, f.hangs + exists, hang := f.exists, f.hangs || op == f.hangOn f.mu.Unlock() - if hangs { + if hang { <-r.Context().Done() return } @@ -259,3 +270,50 @@ func TestNew_DynamoDBDedupeRefusesNoRegion(t *testing.T) { }) } } + +// A reload must not wait behind a tenant's own in-flight DynamoDB call when +// nothing changes for that tenant: Managed.Apply's no-op fast path settles +// under a read lock, so it never contends with a Commit already holding one +// — and, since Go's RWMutex blocks new readers behind a pending writer, a +// concurrent Reserve for the same tenant must also go through, which it +// would not if the reload's Apply took the write lock unconditionally. +func TestReload_DynamoDBDedupeDoesNotWaitOnInFlightCommit(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, true) + fake.setHangOn("BatchWriteItem") + a := newApp(t, cfg, Options{}) + + store := a.dedup.For(tenant.Default) + require.True(t, store.Open()) + claims, err := store.Reserve(context.Background(), []dedupe.Key{eventKey}, time.Minute) + require.NoError(t, err) + require.Equal(t, dedupe.Claimed, claims[0].Status) + + commitCtx, cancelCommit := context.WithCancel(context.Background()) + defer cancelCommit() + commitDone := make(chan error, 1) + go func() { commitDone <- store.Commit(commitCtx, claims, 0) }() + require.Eventually(t, func() bool { return fake.called("BatchWriteItem") }, time.Second, time.Millisecond, + "commit reached the table and is now hanging on it") + + start := time.Now() + a.tenants.Reload("test") + assert.Less(t, time.Since(start), 500*time.Millisecond, + "a reload that changes nothing for this tenant waited on its in-flight commit") + + otherKey := dedupe.Key{Table: eventKey.Table, ID: "concurrent-reserve"} + reserveDone := make(chan error, 1) + go func() { + _, err := store.Reserve(context.Background(), []dedupe.Key{otherKey}, time.Minute) + reserveDone <- err + }() + select { + case err := <-reserveDone: + require.NoError(t, err, "a Reserve for the same tenant, started right after the reload") + case <-time.After(500 * time.Millisecond): + t.Fatal("a concurrent Reserve for the same tenant was blocked") + } + + cancelCommit() + <-commitDone // let the hung call finish (canceled) before the app closes +} From 3235598fc0bb092717a764ebc576c9ce7a56fcde Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:28:05 -0400 Subject: [PATCH 091/108] test(cmd): keep the boot test's broker store in storedir TestRun_BootsAndStopsOnCancel boots the full process through run(), embedded broker included, but still gave it a bare t.TempDir() as data_dir, so its durables' late state writes could fail the removal the same way #442's did. It now uses storedir.New(t); run() returns inside the test body, so the store is removed after the broker has closed. The TestRun_RefusesToBoot cases fail before the broker opens and keep t.TempDir(). The comment on EmbeddedNATS.Close now names the production side of the same gap, #665: the flusher's write can land after Close returns, or never if the process exits first, leaving the durable's previous ack state on disk. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- cmd/wavehouse/main_test.go | 3 ++- internal/mq/embedded.go | 4 +++- 3 files changed, 6 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a4f91f32..0be38b97 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,7 +85,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Tests that start the embedded broker no longer fail removing its store after passing** (`internal/testutil/storedir` (new, + tests), `internal/testutil/testutil.go`, `internal/mq/embedded.go` (comment), `internal/mq/{embedded_test,mqtest/embedded_test}.go`, `internal/ingest/worker_test.go`, `internal/app/{app,roles}_test.go`, `tests/integration/{ingest_outage,query_errors,tenants}_test.go`, `AGENTS.md`, `docs/src/content/docs/development.md`): [#442](https://github.com/Wave-RF/WaveHouse/issues/442). The NATS server writes each durable consumer's state (`obs//o.dat`, through a temporary file renamed into place) from a goroutine that neither `Shutdown` nor `WaitForShutdown` joins, and its consumer store waits for that goroutine at close only when state is still unwritten, for at most 100ms — so a write already under way lands after `EmbeddedNATS.Close` returns, and `t.TempDir`'s one-shot `RemoveAll` met the late entry as `directory not empty`. Under parallel test processes it failed about 4% of the ingest worker tests (78 of 1,800 runs). Every store a test puts on disk now comes from `storedir.New(t)`, whose cleanup — after the broker's `Close` — removes it again whenever a directory was refilled between being read and being removed: each late write adds at most two entries and none once its directory is gone, so the removal ends without a timer (0 of 1,800 under the same load). It replaces two sleep-and-retry copies in the `internal/mq` tests. `TestStartIngestWorker_StopFunc_RespectsShutdownDeadline` also joins the worker its deadline abandons before the broker closes, rather than leaving it to ack on a closed connection. +- **Tests that start the embedded broker no longer fail removing its store after passing** (`internal/testutil/storedir` (new, + tests), `internal/testutil/testutil.go`, `internal/mq/embedded.go` (comment), `internal/mq/{embedded_test,mqtest/embedded_test}.go`, `internal/ingest/worker_test.go`, `internal/app/{app,roles}_test.go`, `cmd/wavehouse/main_test.go`, `tests/integration/{ingest_outage,query_errors,tenants}_test.go`, `AGENTS.md`, `docs/src/content/docs/development.md`): [#442](https://github.com/Wave-RF/WaveHouse/issues/442). The NATS server writes each durable consumer's state (`obs//o.dat`, through a temporary file renamed into place) from a goroutine that neither `Shutdown` nor `WaitForShutdown` joins, and its consumer store waits for that goroutine at close only when state is still unwritten, for at most 100ms — so a write already under way lands after `EmbeddedNATS.Close` returns, and `t.TempDir`'s one-shot `RemoveAll` met the late entry as `directory not empty`. Under parallel test processes it failed about 4% of the ingest worker tests (78 of 1,800 runs). Every store a test puts on disk now comes from `storedir.New(t)`, whose cleanup — after the broker's `Close` — removes it again whenever a directory was refilled between being read and being removed: each late write adds at most two entries and none once its directory is gone, so the removal ends without a timer (0 of 1,800 under the same load). It replaces two sleep-and-retry copies in the `internal/mq` tests. `TestStartIngestWorker_StopFunc_RespectsShutdownDeadline` also joins the worker its deadline abandons before the broker closes, rather than leaving it to ack on a closed connection. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). diff --git a/cmd/wavehouse/main_test.go b/cmd/wavehouse/main_test.go index d0e8404c..ff471474 100644 --- a/cmd/wavehouse/main_test.go +++ b/cmd/wavehouse/main_test.go @@ -20,6 +20,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/config" "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" ) // run reads the whole boot config from the environment here (no config @@ -78,7 +79,7 @@ func seedSettings(t *testing.T) string { func TestRun_BootsAndStopsOnCancel(t *testing.T) { hermeticEnv(t) t.Setenv(config.EnvSettingsDir, seedSettings(t)) - t.Setenv("WH_DATA_DIR", t.TempDir()) + t.Setenv("WH_DATA_DIR", storedir.New(t)) _, port, err := net.SplitHostPort(closedAddr(t)) require.NoError(t, err) t.Setenv("WH_SERVER_PORT", port) diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index c8c7581d..c4928b3c 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -1087,7 +1087,9 @@ func (e *EmbeddedNATS) Close() error { // run()'s remaining defers unwind while JetStream is still tearing down // and the process can exit mid-shutdown (as-if-crashed stream state). // Milliseconds for an in-process server. It does not join a durable's - // state flusher, whose write under way can land after Close returns (#442). + // state flusher: a write under way can land after Close returns, or never + // if the process exits first, leaving the durable's previous ack state on + // disk (#665). e.server.WaitForShutdown() return nil } From 5e1ff2dd1b07448b165646d6a5f211514e141fbf Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:33:30 -0400 Subject: [PATCH 092/108] test(dedupe): replace the sweep-lock timing assertion with structural ones MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits (e52b336c, this PR) raced a Commit against a sweep chunk over 300k tombstones and asserted the slowest attempt stayed under a 100ms wall-clock budget. That budget itself flaked under contention: 313ms measured when the whole internal/dedupe/... race suite ran, since test-unit runs packages in parallel under -race in make ci and a wall-clock bound moves with scheduler load. Replace it with two assertions load cannot move: a new sweepTouchHook (embedded.go, wired into deleteSweepable in sweep.go) tallies how many candidates the locked phase re-reads — asserted <= 1, the actual candidate count in this scenario, proving the locked phase's work is bounded by the candidates the unlocked read found sweepable, not by however many tombstones it silently stepped over inside Pebble to get there. A direct, non-blocking commitMu.TryLock() from within the existing sweepScanHook (fired after the unlocked read, before deleteSweepable's lock) proves the lock was free at that point without racing a goroutine or a clock at all: on the same goroutine that just did the read, TryLock fails instead of blocking if a regression left the lock held, so a pass is a direct proof, not an inference from a race won in time. Verified the new test actually catches the regression it guards against: temporarily wrapped sweepChunk's whole body (including the unlocked read) in e.commitMu.Lock()/Unlock() and dropped deleteSweepable's own lock to avoid a self-deadlock, simulating the pre-fix "sweep under the lock" design — the test failed exactly on the new TryLock assertion ("commitMu was free right after the unlocked read of 300k tombstones"), then reverted; `git diff` on sweep.go before adding the touch hook showed only the hook call, confirming a clean revert. GOTOOLCHAIN=go1.26.6 go test -race -count=20 -run Sweep ./internal/dedupe/ passed 20/20 while a concurrent `go test -race ./internal/...` ran in the background for contention (both processes exited 0). go build ./... and go vet -tags integration ./... are clean. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/embedded.go | 7 ++++- internal/dedupe/sweep.go | 3 ++ internal/dedupe/sweep_test.go | 59 ++++++++++++++++++++--------------- 3 files changed, 42 insertions(+), 27 deletions(-) diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 5eece3f3..7cd60bd5 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -50,9 +50,14 @@ type Embedded struct { readHook func() error // sweepScanHook and sweepDeleteHook, when set, run in a sweep chunk: // between its unlocked read and its re-read, and between its re-read and - // its delete. A test races a Commit into each gap. + // its delete. A test races a Commit into each gap. sweepTouchHook, when + // set, runs once per candidate deleteSweepable re-reads under commitMu — + // a test tallies calls to pin that the locked phase's work is bounded by + // the candidate count, not by however many keys the unlocked read + // stepped over (silently, inside Pebble) to find them. sweepScanHook func() sweepDeleteHook func() + sweepTouchHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index f1c4ca18..d5a49149 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -162,6 +162,9 @@ func (e *Embedded) deleteSweepable(db *pebble.DB, candidates [][]byte) (expired, b := db.NewBatch() defer func() { _ = b.Close() }() for _, k := range candidates { + if e.sweepTouchHook != nil { + e.sweepTouchHook() + } val, closer, err := db.Get(k) if errors.Is(err, pebble.ErrNotFound) { continue diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index 0c7ff46e..493e3745 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -161,45 +161,52 @@ func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { // A chunk that starts over a long run of tombstones, as a tenant's version-0 // block leaves until Pebble compacts it, reads through the run without the -// lock, so a Commit racing it waits for the chunk's re-reads and deletes -// alone. Sized for the race detector, which the unit suite runs under: there, -// a chunk holding the lock across this run kept a Commit waiting ~250 ms. +// lock: deleteSweepable's locked phase only ever touches the candidates the +// unlocked read found sweepable (one here), never the tombstones stepped +// over to find them, and commitMu is free the instant that read returns. +// +// This used to race a Commit against the sweep and assert its slowest +// attempt stayed under 100ms — a wall-clock budget the race detector alone +// could push past (250ms measured once), and one that moves further under +// make ci's parallel test-unit packages (313ms measured there once, this +// PR's e52b336c: the test that introduced it flaked in its own author's +// verification). A count and a direct, non-blocking TryLock don't move with +// scheduler contention, so they replace it. func TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits(t *testing.T) { t.Parallel() e := NewEmbedded(t.TempDir()) - m := switchedOn(t, e, "acme") + switchedOn(t, e, "acme") b := e.db.NewBatch() for i := range 300_000 { require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%06d", i), nil)) } - // A version-0 key after the run, so the chunk has one to delete. + // A version-0 key after the run, so the chunk has exactly one candidate. require.NoError(t, b.Set([]byte("acme\x01"), make([]byte, 8), nil)) require.NoError(t, b.Commit(pebble.NoSync)) require.NoError(t, e.db.Flush()) - type swept struct { - res sweepResult - err error - } - done := make(chan swept, 1) - go func() { - res, err := e.sweep(context.Background(), e.db) - done <- swept{res, err} - }() - var slowest time.Duration - for i := 0; ; i++ { - select { - case s := <-done: - require.NoError(t, s.err) - require.Equal(t, sweepResult{Version0: 1}, s.res) - assert.Less(t, slowest, 100*time.Millisecond, "the slowest Commit racing the sweep") - return - default: + var touched int + e.sweepTouchHook = func() { touched++ } + var unlockedAfterScan bool + e.sweepScanHook = func() { + // Fires after the unlocked read, before deleteSweepable takes + // commitMu. TryLock needs no racing goroutine and no clock: on the + // same goroutine that just did the read, it fails instead of + // blocking if commitMu is already held — which a regression moving + // the read under the lock would leave it, self-deadlock included — + // so success here is a direct proof the read ran unlocked, not an + // inference from a race won in time. + if e.commitMu.TryLock() { + e.commitMu.Unlock() + unlockedAfterScan = true } - start := time.Now() - commitIDs(t, m, 0, fmt.Sprintf("c%d", i)) - slowest = max(slowest, time.Since(start)) } + + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Version0: 1}, res) + assert.True(t, unlockedAfterScan, "commitMu was free right after the unlocked read of 300k tombstones") + assert.LessOrEqual(t, touched, 1, "the locked phase touches only the sweepable candidates (1 here), not the tombstones the read stepped over to find them") } // A retention is honoured on read before any sweep has run: the key is a From b16c193ed395c1f02b6b2b5c82e26f5a29e36d42 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:38:22 -0400 Subject: [PATCH 093/108] test(app): drop the concurrent-Reserve half of the reload/commit test MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The Reserve only started after Reload had already returned, so no writer was ever pending when it ran, and it could not fail for the reason its comment gave — with the fast path removed only the reload-time assertion actually failed. No clean sync point exists into "Reload is inside the dedupe hook" without instrumenting production code for the test alone, so drop the block and the sentence claiming it; the reload-time assertion already pins the regression on its own. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/app/dedupe_dynamodb_test.go | 19 ++----------------- 1 file changed, 2 insertions(+), 17 deletions(-) diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index 71bf294c..310d1140 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -273,10 +273,8 @@ func TestNew_DynamoDBDedupeRefusesNoRegion(t *testing.T) { // A reload must not wait behind a tenant's own in-flight DynamoDB call when // nothing changes for that tenant: Managed.Apply's no-op fast path settles -// under a read lock, so it never contends with a Commit already holding one -// — and, since Go's RWMutex blocks new readers behind a pending writer, a -// concurrent Reserve for the same tenant must also go through, which it -// would not if the reload's Apply took the write lock unconditionally. +// under a read lock alone, so it never contends with a Commit already +// holding one and returns long before the commit does. func TestReload_DynamoDBDedupeDoesNotWaitOnInFlightCommit(t *testing.T) { cfg := testConfig(t, writeSettings(t, dedupeOn)) fake := dynamoConfig(t, cfg, true) @@ -301,19 +299,6 @@ func TestReload_DynamoDBDedupeDoesNotWaitOnInFlightCommit(t *testing.T) { assert.Less(t, time.Since(start), 500*time.Millisecond, "a reload that changes nothing for this tenant waited on its in-flight commit") - otherKey := dedupe.Key{Table: eventKey.Table, ID: "concurrent-reserve"} - reserveDone := make(chan error, 1) - go func() { - _, err := store.Reserve(context.Background(), []dedupe.Key{otherKey}, time.Minute) - reserveDone <- err - }() - select { - case err := <-reserveDone: - require.NoError(t, err, "a Reserve for the same tenant, started right after the reload") - case <-time.After(500 * time.Millisecond): - t.Fatal("a concurrent Reserve for the same tenant was blocked") - } - cancelCommit() <-commitDone // let the hung call finish (canceled) before the app closes } From 674746a78ab6de0049ea23558361af61c2178acd Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:38:42 -0400 Subject: [PATCH 094/108] docs(dedupe): the idle-pool floor, and what a reload's wait covers configuration.mdx's dedupe.reserve_concurrency row now says the DynamoDB idle-connection pool never goes below the SDK's own default (10), matching newHTTPClient's max(...) floor. configuration.mdx, deployment.md and CHANGELOG.md said switching a tenant's dedupe off was the one case a reload waits on. It is not the only one: a reload that removes or rejects a tenant whose store was open closes it the same way (Stores.Retain -> Managed.Close -> Apply(false)), taking the same write lock and waiting on the same in-flight calls. Reworded all three to match architecture.md's existing "on a genuine flip" framing, which already covered both cases. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 4 ++-- docs/src/content/docs/deployment.md | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 747fbd8f..4ab522f6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; switching a tenant's dedupe off is the exception, waiting for that tenant's in-flight calls to finish before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index a1286b53..bd1781d0 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -57,11 +57,11 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, the lease plus its own ceiling to the next whole second plus one more second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59s`: a client that obeys `Retry-After` after a publish whose outcome it never learned can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match. Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match, never fewer than the SDK's own default (10). Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. Switching a tenant's dedupe off is the exception — it waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish before the store closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index e9aca371..ddf37795 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; switching a tenant's dedupe off waits for that tenant's in-flight calls to finish before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. From 51716a29bb6b9c8f3bc7b76045a300b9e612679d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:42:28 -0400 Subject: [PATCH 095/108] test(api): pin the ERROR/DEBUG split on a Reserve failure by exact msg MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError's comment claimed TestIngest_Dedup_ReserveError pins the real-backend-failure case staying ERROR — it doesn't; that test only asserts status, body and publish count, never the log. Two mutations passed the whole internal/api suite as a result: `if ctx.Err() != nil` -> `if true` (collapses every Reserve failure onto the DEBUG branch), and the client-gone line's DebugContext -> WarnContext (both leave a bare `"level":"DEBUG"` check satisfied by an unrelated "debug: span started for ingest" line the package logs on every request). Replace the single-case test with one non-parallel, two-case TestIngest_Dedup_ReserveError_LogLevel: a live context asserts the line `"level":"ERROR","msg":"dedupe reserve failed"`; a cancelled one asserts `"level":"DEBUG","msg":"dedupe reserve failed: request context ended"`. Matching level and msg as one adjacent substring (the exact order slog's JSON handler emits them in) ties the level to the specific line rather than to the buffer as a whole, so neither mutation above can hide behind the unrelated DEBUG line. Verified both mutations now fail the new test (live-context case fails under `if true`; context-ended case fails under WarnContext), then reverted each — `git diff` on ingest.go showed no residue before the two doc fixes below were applied. Also, [MAY]: requestAbort.RetryAfter's inline comment repeated a narrower, now-stale cause list (missing the dedupe-store-unavailable 503) that duplicates and drifts from the requestAbort doc comment above it, which already owns that list — trimmed to the field's own job. mq.ErrUnavailable's doc said the API answers it with "a short Retry-After"; publishFailed sends the full dedupe lease (rounded up to whole seconds) when the failing record held a claim, and only falls back to a flat few seconds otherwise — reworded to say so. GOTOOLCHAIN=go1.26.6 go test -race ./internal/api/... ./internal/mq/... passes. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest.go | 2 +- internal/api/ingest_test.go | 66 ++++++++++++++++++++++++++----------- internal/mq/mq.go | 7 ++-- 3 files changed, 52 insertions(+), 23 deletions(-) diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 2007bb7b..c85d52c3 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -158,7 +158,7 @@ type recordReject struct { type requestAbort struct { Status int Message string - RetryAfter string // non-empty → emit a Retry-After header (503: backpressure, an unavailable broker, or an id another request holds) + RetryAfter string // non-empty → emit a Retry-After header } func (h *IngestHandler) Handle(w http.ResponseWriter, r *http.Request) { diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index 0e1b6d52..034a329c 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -2987,29 +2987,55 @@ func TestIngest_Dedup_ReserveError(t *testing.T) { assert.Empty(t, pub.Published()) } -// A Reserve error caused by the request's own context ending (the client -// gone, or its deadline past) is not a backend failure and must not log at -// ERROR — an operator paging on ERROR logs would otherwise be woken by -// clients that simply went away. TestIngest_Dedup_ReserveError above pins the -// real-backend-failure case, which stays ERROR. -func TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError(t *testing.T) { - buf := logtest.Capture(t, slog.LevelDebug) - pub := &testutil.MockPublisher{} - dedup := testutil.NewMockDeduplicator() - dedup.Err = errors.New("backend down") - h := dedupHandler(t, pub, dedup, false) +// A Reserve error's log level depends on why it failed. A real backend +// failure (a live request context) stays ERROR, so an operator is paged. One +// caused by the request's own context ending (the client gone, or its +// deadline past) is not a backend problem and must log at DEBUG instead — an +// operator paging on ERROR logs would otherwise be woken by clients that +// simply went away. Not t.Parallel: it captures the process-wide default +// logger (logtest.Capture). Matched on the exact "level":"…","msg":"…" pair +// slog's JSON handler emits adjacently, not on the level alone — the package +// also logs an unrelated "debug: span started for ingest" line per request, +// which satisfies a bare `"level":"DEBUG"` check whether or not the Reserve +// line itself is DEBUG. +func TestIngest_Dedup_ReserveError_LogLevel(t *testing.T) { + tests := []struct { + name string + cancelContext bool + wantLine string + }{ + { + "live context: a real backend failure pages at ERROR", false, + `"level":"ERROR","msg":"dedupe reserve failed"`, + }, + { + "context ended: a client gone must not page, logs at DEBUG", true, + `"level":"DEBUG","msg":"dedupe reserve failed: request context ended"`, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + buf := logtest.Capture(t, slog.LevelDebug) + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.Err = errors.New("backend down") + h := dedupHandler(t, pub, dedup, false) - ctx, cancel := context.WithCancel(context.Background()) - cancel() - req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}).WithContext(ctx) + ctx := context.Background() + if tt.cancelContext { + var cancel context.CancelFunc + ctx, cancel = context.WithCancel(ctx) + cancel() + } + req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}).WithContext(ctx) - w := httptest.NewRecorder() - h.Handle(w, withTenant(req)) + w := httptest.NewRecorder() + h.Handle(w, withTenant(req)) - assert.Contains(t, buf.String(), "dedupe reserve failed", "still logged, just not at ERROR") - assert.NotContains(t, buf.String(), `"level":"ERROR"`, "a client-gone Reserve error must not page an operator") - assert.Contains(t, buf.String(), `"level":"DEBUG"`) - assert.Empty(t, pub.Published()) + assert.Contains(t, buf.String(), tt.wantLine) + assert.Empty(t, pub.Published()) + }) + } } // #370: an explicit null id is a missing id — rejected under require_id, diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 6415e481..f88bcfd5 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -192,8 +192,11 @@ var ErrQueueFull = errors.New("ingest queue is full") // ErrUnavailable is returned when the broker cannot be reached or does not // answer in time — a transient failure, not a refusal, that the API turns -// into a 503 with a short Retry-After. Only a backend whose broker is out of -// process returns it; the embedded one's publish failures are plain errors. +// into a 503. Retry-After is the dedupe lease, rounded up to whole seconds, +// when the record held a claim (so an obedient client waits out the window +// instead of retrying straight into it), else a flat few seconds. Only a +// backend whose broker is out of process returns it; the embedded one's +// publish failures are plain errors. var ErrUnavailable = errors.New("message queue unavailable") // Publisher appends events to the ingest queue. From c4147ecefefb0686057b16ab2a20a39c85b1a538 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:53:41 -0400 Subject: [PATCH 096/108] test(dedupe): pin sweepCandidates' read as unlocked from inside its loop MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits claimed more than it checked. sweepTouchHook fired once per element of `candidates`, and the fixture has exactly one visible key, so `touched <= 1` could never fail: a full keyspace walk added inside the locked phase would still pass (R4). The TryLock ran in sweepScanHook, which fires AFTER sweepCandidates returns, so a regression that read under commitMu and released the lock just before that hook still passed (R3) — only "the whole chunk under one lock" was actually caught (R5). The 300k-tombstone fixture changed no verdict either way (a handful of tombstones behaves identically) while costing 1.65s alone / 3.1s under package load in a 15s-timeout package, and the comment recorded this test's own history (citing a commit that a squash would erase) instead of what it asserts. Drop sweepTouchHook (field, wiring, assertion). To pin "the read runs unlocked" for real, sweepCandidates now takes an onKey hook called once per key from inside its own loop, before evaluating it — wired through sweepChunk as e.sweepReadHook. Renamed TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits to TestEmbedded_SweepReadRunsUnlocked: TryLock/Unlock from inside that hook, while the read is still running, so a lock held anywhere during the read is caught in the act rather than inferred from whether it was released before some later checkpoint. The fixture shrinks to a handful of tombstones ahead of the one live key, since the verdict never depended on the count. Comment cut to the one thing the test asserts; the old wall-clock/hook history belongs here instead. Verified by mutation, each applied then reverted (`git diff internal/dedupe/sweep.go` clean before the real change was made): - R3 (only the read under the lock, released right after): wrapped just the sweepCandidates call in sweepChunk with commitMu.Lock()/Unlock() — TestEmbedded_SweepReadRunsUnlocked failed ("commitMu must be free while sweepCandidates' read is running"). - R5 (the whole chunk under one lock): wrapped sweepChunk's body in commitMu.Lock()/defer Unlock() and dropped deleteSweepable's own lock to avoid a self-deadlock — ran ONLY the target test (not the package: TestEmbedded_SweepNeverDeletesACommitLandingMidChunk self-deadlocks under this mutation, racing a Commit against a sweep that never releases the lock) — failed with the same assertion. GOTOOLCHAIN=go1.26.6 go test -race -count=5 -run Sweep ./internal/dedupe/ passes (5/5, including TestEmbedded_SweepReadRunsUnlocked and every other Sweep-prefixed test); the full package (go test -race -count=1 ./internal/dedupe/...) passes too, now in ~5s rather than the prior fixture's ~57s at -count=20. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/embedded.go | 11 ++++---- internal/dedupe/sweep.go | 13 +++++----- internal/dedupe/sweep_test.go | 47 ++++++++++++----------------------- 3 files changed, 28 insertions(+), 43 deletions(-) diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 7cd60bd5..227564b7 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -50,14 +50,13 @@ type Embedded struct { readHook func() error // sweepScanHook and sweepDeleteHook, when set, run in a sweep chunk: // between its unlocked read and its re-read, and between its re-read and - // its delete. A test races a Commit into each gap. sweepTouchHook, when - // set, runs once per candidate deleteSweepable re-reads under commitMu — - // a test tallies calls to pin that the locked phase's work is bounded by - // the candidate count, not by however many keys the unlocked read - // stepped over (silently, inside Pebble) to find them. + // its delete. A test races a Commit into each gap. sweepReadHook, when + // set, runs once per key sweepCandidates' unlocked read visits, from + // inside its loop — a test TryLocks commitMu there to prove the read + // itself never holds it, not just the instant after it returns. sweepScanHook func() sweepDeleteHook func() - sweepTouchHook func() + sweepReadHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index d5a49149..25759f39 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -104,7 +104,7 @@ func (e *Embedded) sweep(ctx context.Context, db *pebble.DB) (sweepResult, error // read and a run of them left by an earlier pass would otherwise hold every // Commit for its whole length. func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, res *sweepResult) ([]byte, error) { - candidates, next, err := sweepCandidates(db, from, e.now()) + candidates, next, err := sweepCandidates(db, from, e.now(), e.sweepReadHook) if err != nil || len(candidates) == 0 { return next, err } @@ -128,14 +128,18 @@ func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, r // sweepCandidates reads the next sweepChunk keys, starting at from, and // returns those sweepable at now and where the next chunk starts (nil at the -// end). -func sweepCandidates(db *pebble.DB, from []byte, now time.Time) (candidates [][]byte, next []byte, err error) { +// end). onKey, when non-nil, runs once per key visited, before it is +// evaluated — a test hook proving this read holds no lock while it runs. +func sweepCandidates(db *pebble.DB, from []byte, now time.Time, onKey func()) (candidates [][]byte, next []byte, err error) { it, err := db.NewIter(&pebble.IterOptions{LowerBound: from}) if err != nil { return nil, nil, fmt.Errorf("dedupe sweep: %w", err) } seen := 0 for valid := it.First(); valid; valid = it.Next() { + if onKey != nil { + onKey() + } if seen == sweepChunk { next = bytes.Clone(it.Key()) break @@ -162,9 +166,6 @@ func (e *Embedded) deleteSweepable(db *pebble.DB, candidates [][]byte) (expired, b := db.NewBatch() defer func() { _ = b.Close() }() for _, k := range candidates { - if e.sweepTouchHook != nil { - e.sweepTouchHook() - } val, closer, err := db.Get(k) if errors.Is(err, pebble.ErrNotFound) { continue diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index 493e3745..a0255e39 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -159,54 +159,39 @@ func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { assert.True(t, dup, "the commit made mid-chunk survived the sweep") } -// A chunk that starts over a long run of tombstones, as a tenant's version-0 -// block leaves until Pebble compacts it, reads through the run without the -// lock: deleteSweepable's locked phase only ever touches the candidates the -// unlocked read found sweepable (one here), never the tombstones stepped -// over to find them, and commitMu is free the instant that read returns. -// -// This used to race a Commit against the sweep and assert its slowest -// attempt stayed under 100ms — a wall-clock budget the race detector alone -// could push past (250ms measured once), and one that moves further under -// make ci's parallel test-unit packages (313ms measured there once, this -// PR's e52b336c: the test that introduced it flaked in its own author's -// verification). A count and a direct, non-blocking TryLock don't move with -// scheduler contention, so they replace it. -func TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits(t *testing.T) { +// sweepCandidates' read never holds commitMu, over a fixture with a few +// tombstones ahead of the one live key it finds sweepable. +func TestEmbedded_SweepReadRunsUnlocked(t *testing.T) { t.Parallel() e := NewEmbedded(t.TempDir()) switchedOn(t, e, "acme") b := e.db.NewBatch() - for i := range 300_000 { - require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%06d", i), nil)) + for i := range 4 { + require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%02d", i), nil)) } - // A version-0 key after the run, so the chunk has exactly one candidate. require.NoError(t, b.Set([]byte("acme\x01"), make([]byte, 8), nil)) require.NoError(t, b.Commit(pebble.NoSync)) require.NoError(t, e.db.Flush()) - var touched int - e.sweepTouchHook = func() { touched++ } - var unlockedAfterScan bool - e.sweepScanHook = func() { - // Fires after the unlocked read, before deleteSweepable takes - // commitMu. TryLock needs no racing goroutine and no clock: on the - // same goroutine that just did the read, it fails instead of - // blocking if commitMu is already held — which a regression moving - // the read under the lock would leave it, self-deadlock included — - // so success here is a direct proof the read ran unlocked, not an - // inference from a race won in time. + var sawLocked bool + e.sweepReadHook = func() { + // TryLock from inside the still-running read needs no racing + // goroutine and no clock: on the same goroutine doing the read, it + // fails instead of blocking if commitMu is already held, so a + // regression that reads under the lock is caught while the read is + // still in progress — not inferred from a race won in time, and not + // missable by a lock released just before some later checkpoint. if e.commitMu.TryLock() { e.commitMu.Unlock() - unlockedAfterScan = true + } else { + sawLocked = true } } res, err := e.sweep(context.Background(), e.db) require.NoError(t, err) assert.Equal(t, sweepResult{Version0: 1}, res) - assert.True(t, unlockedAfterScan, "commitMu was free right after the unlocked read of 300k tombstones") - assert.LessOrEqual(t, touched, 1, "the locked phase touches only the sweepable candidates (1 here), not the tombstones the read stepped over to find them") + assert.False(t, sawLocked, "commitMu must be free while sweepCandidates' read is running") } // A retention is honoured on read before any sweep has run: the key is a From 21380681f230f02e72f13b6d44d67c080e46a784 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:55:04 -0400 Subject: [PATCH 097/108] test: keep this branch's two broker stores in storedir TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat and the api package's realPipeline opened the embedded broker on a bare t.TempDir(). Both replay through a disk-backed consumer, so they were exposed to the late consumer-state write (#442) that storedir absorbs. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest_window_test.go | 3 ++- internal/mq/embedded_test.go | 2 +- 2 files changed, 3 insertions(+), 2 deletions(-) diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index ba350eed..73bb83bb 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -15,6 +15,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/settings" "github.com/Wave-RF/WaveHouse/internal/testutil" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" ) @@ -289,7 +290,7 @@ func (p *faultyPublisher) Publish(ctx context.Context, topic mq.Topic, data []by // in the tenant's queue. func realPipeline(t *testing.T, fail func(call int) (bool, error)) (*IngestHandler, func() int) { t.Helper() - broker, err := mq.NewEmbedded(t.TempDir()) + broker, err := mq.NewEmbedded(storedir.New(t)) require.NoError(t, err) t.Cleanup(func() { _ = broker.Close() }) require.NoError(t, broker.SetMaxBytes(t.Context(), testStore.Tenant(), 64<<20)) diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index a2fced64..e1d28281 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -151,7 +151,7 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // path, where takeStock itself must not mistake a stale window for one // already at budget. func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { - e := openEmbedded(t, t.TempDir()) + e := openEmbedded(t, storedir.New(t)) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() // Explicit rather than the server's default, which happens to match today. From a76bbb9d86a6188b07413b9c6e64afe9d51b85e1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:00:32 -0400 Subject: [PATCH 098/108] test(dedupe): assert the sweep's read hook ran TestEmbedded_SweepReadRunsUnlocked asserted only that the hook never saw commitMu held, which also holds if the hook never fires: passing nil for the hook left it green. It now counts the visits and requires one; with the hook unwired it fails. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/sweep_test.go | 10 ++++------ 1 file changed, 4 insertions(+), 6 deletions(-) diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index a0255e39..ee87088e 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -173,14 +173,11 @@ func TestEmbedded_SweepReadRunsUnlocked(t *testing.T) { require.NoError(t, b.Commit(pebble.NoSync)) require.NoError(t, e.db.Flush()) + var visits int var sawLocked bool e.sweepReadHook = func() { - // TryLock from inside the still-running read needs no racing - // goroutine and no clock: on the same goroutine doing the read, it - // fails instead of blocking if commitMu is already held, so a - // regression that reads under the lock is caught while the read is - // still in progress — not inferred from a race won in time, and not - // missable by a lock released just before some later checkpoint. + visits++ + // Non-blocking, on the reading goroutine: fails if the read holds commitMu. if e.commitMu.TryLock() { e.commitMu.Unlock() } else { @@ -191,6 +188,7 @@ func TestEmbedded_SweepReadRunsUnlocked(t *testing.T) { res, err := e.sweep(context.Background(), e.db) require.NoError(t, err) assert.Equal(t, sweepResult{Version0: 1}, res) + assert.Positive(t, visits, "the read hook ran") assert.False(t, sawLocked, "commitMu must be free while sweepCandidates' read is running") } From 7f38b9b350924f4eede8e3f45951edb72c574537 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:00:41 -0400 Subject: [PATCH 099/108] fix(app): move DynamoDB dedupe wiring out of wire.go for the e2e gate make ci's e2e coverage measured 59.7% against a 60% floor: the e2e binary always boots with dedupe.backend: pebble, so wireDynamoDedupe, errDynamoUnchecked and the table-check retry component never ran there, same shape as #628's internal/dedupe/dynamodb.go exclusion. Pure move, no behavior change: wireDynamoDedupe and errDynamoUnchecked move verbatim into the new internal/app/wire_dynamodb.go; wireDedupe's switch stays in wire.go untouched. Adds a matching e2e-only exclusion for the new file in .testcoverage.yml, next to dynamodb.go's, so unit and integration keep covering it and the merged total still counts it -- wire.go itself stays out of the exclude list. Updates architecture.md's wire.go bullet (the dynamodb case now points at the new file's own bullet) and the CHANGELOG's file-provenance list for the dedupe.backend: dynamodb entry. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- .testcoverage.yml | 6 ++ CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 +- internal/app/wire.go | 115 +----------------------- internal/app/wire_dynamodb.go | 125 ++++++++++++++++++++++++++ 5 files changed, 139 insertions(+), 113 deletions(-) create mode 100644 internal/app/wire_dynamodb.go diff --git a/.testcoverage.yml b/.testcoverage.yml index 77b365e9..afb04f49 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -84,3 +84,9 @@ exclude: # (fake API) and integration (dynamodb-local) suites cover it, and the # merged total still counts it. - ^internal/dedupe/dynamodb\.go$ + # wireDynamoDedupe and its retry component (internal/app/wire_dynamodb.go): + # same reason as dynamodb.go above — the e2e binary never selects + # dedupe.backend: dynamodb, so this file measured 0% there and pulled + # e2e to 59.7%. The unit and integration suites cover it, and the + # merged total still counts it. + - ^internal/app/wire_dynamodb\.go$ diff --git a/CHANGELOG.md b/CHANGELOG.md index 574c8f86..5d405dc6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6f18bb5a..579f5a21 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -93,7 +93,9 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. `wireDedupe`'s `dynamodb` case is `wireDynamoDedupe`, in wire_dynamodb.go (below). `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. + +- **wire_dynamodb.go** — `wireDedupe`'s `dynamodb` case, split out of wire.go so the e2e suite's coverage exclude for it (the e2e binary always runs Pebble dedupe, never DynamoDB) doesn't have to blanket wire.go itself: builds the same `dedupe.Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. ### `stream/` — SSE keepalive & fan-out diff --git a/internal/app/wire.go b/internal/app/wire.go index a65d1bf8..dd48fa07 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -532,117 +532,10 @@ func (a *App) wirePebbleDedupe() error { return nil } -// errDynamoUnchecked is a store's open before the first table check has run. -var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") - -// wireDynamoDedupe builds the dedupe stores over one DynamoDB table that -// every tenant and every process shares (dedupe.Dynamo), so a tenant's store -// opens for free once the table has passed its check. Boot checks it (after -// creating it, with create_table on dynamodb-local) whether or not any tenant -// has dedupe on, and never creates it otherwise. A table that fails the check -// follows the registry's rule for the shape, as Pebble's instance does: a -// flat directory refuses boot; a nested one boots with every switched-on -// store closed, so its ingest fails closed. Unlike a local disk, a remote -// table's failure is usually brief (a throttle, credentials not yet issued -// mid-rollout), and a nested directory has no watcher to reload it, so the -// check is then retried in the background, with backoff, until it passes. -// The check is network I/O, so the AfterAdopt hook never runs it: the hook -// holds the lock that serializes reloads. It applies every store against the -// last check's result and wakes the retry, so a reload still retries at once. -func (a *App) wireDynamoDedupe(ctx context.Context) error { - c := a.cfg.Dedupe.DynamoDB - d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ - Table: c.Table, Region: c.Region, Endpoint: c.Endpoint, - Timeout: c.Timeout, MaxAttempts: c.MaxAttempts, RetryMode: c.RetryMode, - ReserveConcurrency: a.cfg.Dedupe.ReserveConcurrency, - }) - if err != nil { - return err - } - var mu sync.Mutex - state := errDynamoUnchecked // nil once the table has passed, for good - ready := func() error { - mu.Lock() - defer mu.Unlock() - return state - } - // check is only ever run by boot, then by the retry loop, one at a time. - check := func(ctx context.Context) error { - var err error - if c.CreateTable { - err = d.CreateTable(ctx) - } - if err == nil { - err = d.Check(ctx) - } - mu.Lock() - defer mu.Unlock() - if state != nil { - state = err - } - return state - } - stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) - a.dedup = stores - a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) - var reconciling sync.Mutex // the hook and the retry loop both apply - apply := func() { - reconciling.Lock() - defer reconciling.Unlock() - if err := stores.Retain(a.served); err != nil { - slog.Error("dedupe store close failed", "error", err) - } - for id, store := range a.tenants.All() { - m := stores.For(id) - enabled := store.DedupeEnabled() - wasOpen := m.Open() - // The one failure an open has is the check's, logged where it ran. - _ = m.Apply(enabled) - if m.Open() != wasOpen { - slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) - } - } - } - retry := make(chan struct{}, 1) - a.tenants.AfterAdopt(func([]tenant.ID) { - apply() - if ready() != nil { - select { - case retry <- struct{}{}: - default: // a retry is already due - } - } - }) - if err := check(ctx); err != nil { - if !a.tenants.Nested() { - return fmt.Errorf("dedupe open: %w", err) - } - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", - "table", c.Table, "error", err) - a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { - for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { - select { - case <-ctx.Done(): - return nil - case <-time.After(wait): - case <-retry: - } - if err := check(ctx); err != nil { - if ctx.Err() == nil { - slog.Error("dedupe: dynamodb table check failed again; ingest with dedupe on still fails closed", - "table", c.Table, "error", err) - } - continue - } - slog.Info("dedupe: dynamodb table check passed", "table", c.Table) - apply() - } - return nil - }}) - } - apply() - return nil -} +// wireDynamoDedupe (dedupe.backend: dynamodb) lives in wire_dynamodb.go, +// excluded from the e2e coverage gate alongside internal/dedupe/dynamodb.go +// (see .testcoverage.yml): the e2e binary always runs Pebble dedupe, so +// nothing there exercises it. wireDedupe above still switches on it. // wireMQ starts the MQ — the one place the implementation is chosen; // everything after it sees mq.Broker. diff --git a/internal/app/wire_dynamodb.go b/internal/app/wire_dynamodb.go new file mode 100644 index 00000000..420111b5 --- /dev/null +++ b/internal/app/wire_dynamodb.go @@ -0,0 +1,125 @@ +package app + +import ( + "context" + "errors" + "fmt" + "log/slog" + "sync" + "time" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// errDynamoUnchecked is a store's open before the first table check has run. +var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") + +// wireDynamoDedupe builds the dedupe stores over one DynamoDB table that +// every tenant and every process shares (dedupe.Dynamo), so a tenant's store +// opens for free once the table has passed its check. Boot checks it (after +// creating it, with create_table on dynamodb-local) whether or not any tenant +// has dedupe on, and never creates it otherwise. A table that fails the check +// follows the registry's rule for the shape, as Pebble's instance does: a +// flat directory refuses boot; a nested one boots with every switched-on +// store closed, so its ingest fails closed. Unlike a local disk, a remote +// table's failure is usually brief (a throttle, credentials not yet issued +// mid-rollout), and a nested directory has no watcher to reload it, so the +// check is then retried in the background, with backoff, until it passes. +// The check is network I/O, so the AfterAdopt hook never runs it: the hook +// holds the lock that serializes reloads. It applies every store against the +// last check's result and wakes the retry, so a reload still retries at once. +func (a *App) wireDynamoDedupe(ctx context.Context) error { + c := a.cfg.Dedupe.DynamoDB + d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ + Table: c.Table, Region: c.Region, Endpoint: c.Endpoint, + Timeout: c.Timeout, MaxAttempts: c.MaxAttempts, RetryMode: c.RetryMode, + ReserveConcurrency: a.cfg.Dedupe.ReserveConcurrency, + }) + if err != nil { + return err + } + var mu sync.Mutex + state := errDynamoUnchecked // nil once the table has passed, for good + ready := func() error { + mu.Lock() + defer mu.Unlock() + return state + } + // check is only ever run by boot, then by the retry loop, one at a time. + check := func(ctx context.Context) error { + var err error + if c.CreateTable { + err = d.CreateTable(ctx) + } + if err == nil { + err = d.Check(ctx) + } + mu.Lock() + defer mu.Unlock() + if state != nil { + state = err + } + return state + } + stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) + a.dedup = stores + a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) + var reconciling sync.Mutex // the hook and the retry loop both apply + apply := func() { + reconciling.Lock() + defer reconciling.Unlock() + if err := stores.Retain(a.served); err != nil { + slog.Error("dedupe store close failed", "error", err) + } + for id, store := range a.tenants.All() { + m := stores.For(id) + enabled := store.DedupeEnabled() + wasOpen := m.Open() + // The one failure an open has is the check's, logged where it ran. + _ = m.Apply(enabled) + if m.Open() != wasOpen { + slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) + } + } + } + retry := make(chan struct{}, 1) + a.tenants.AfterAdopt(func([]tenant.ID) { + apply() + if ready() != nil { + select { + case retry <- struct{}{}: + default: // a retry is already due + } + } + }) + if err := check(ctx); err != nil { + if !a.tenants.Nested() { + return fmt.Errorf("dedupe open: %w", err) + } + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", + "table", c.Table, "error", err) + a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { + for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { + select { + case <-ctx.Done(): + return nil + case <-time.After(wait): + case <-retry: + } + if err := check(ctx); err != nil { + if ctx.Err() == nil { + slog.Error("dedupe: dynamodb table check failed again; ingest with dedupe on still fails closed", + "table", c.Table, "error", err) + } + continue + } + slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + apply() + } + return nil + }}) + } + apply() + return nil +} From dc9a331f324cd54f9a64dc5f4b61ff3cad047dcd Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:30:43 -0400 Subject: [PATCH 100/108] fix(dedupe): a caller's cancel no longer strands a DynamoDB claim Puts already sent, and the release after them, run on context.WithoutCancel(ctx), each call still bounded by its Timeout. A client that disconnects mid-Reserve no longer leaves an abandoned put to land after its release and hold the id InFlight for the lease; a Reserve whose caller cancelled after every put answered also releases them. The breaker's exemption for cancelled puts is gone: a sent put can no longer be cancelled by its caller, so its answer is always the table's. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 37 ++++++----- internal/dedupe/dynamodb_test.go | 90 ++++++++++++++++++++------- 4 files changed, 89 insertions(+), 42 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7e299639..b1f5e361 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` leaves nothing claimed: puts already sent run to their answer and are then released, so its retry is not answered `InFlight` for the lease. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7935f25c..20e53dcc 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -137,7 +137,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt-123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped and joined (`keyenc.AppendJoin`) by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with at least as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with at least as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, or the caller cancels (a client disconnecting mid-request), the puts not yet sent are skipped and every put that may have landed is released by its token. A put already sent runs to its answer on a context the caller's cancellation does not reach, so it answers before that undo; only its own call deadline can cut it off, and DynamoDB may then apply it after its release, holding its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index cc3321f7..e9e8cd47 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -303,10 +303,7 @@ func (d *Dynamo) call(ctx context.Context, op string, do func(context.Context) e start := time.Now() err := classify(op, do(ctx)) d.metrics.record(ctx, op, time.Since(start), err) - // A request cancelled because its caller went away (a client - // disconnecting mid-Reserve) says nothing about the table, and must not - // reset the breaker's count. - if op == opReserve && !errors.Is(err, context.Canceled) { + if op == opReserve { d.breaker.record(err) } return err @@ -321,9 +318,10 @@ type dynamoStore struct { // Reserve puts every key's pending item in parallel, each conditional on no // live item holding the key. A failed condition hands back the live item, // whose state says Duplicate or InFlight without a read, or Claimed when its -// token is the put's own. On any error it releases, by token, every put that -// may have landed; what that undo misses (below) holds its key InFlight -// until the lease ends, as a crashed request's claim does. +// token is the put's own. On any error, the caller's cancellation included, +// it releases, by token, every put that may have landed; what that undo +// misses (below) holds its key InFlight until the lease ends, as a crashed +// request's claim does. func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { if len(keys) == 0 { return []Claim{}, nil @@ -338,12 +336,13 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati claims := make([]Claim, len(keys)) tried := make([]Claim, len(keys)) sent := make([]bool, len(keys)) - // The first failure skips the puts not yet sent: the Reserve fails - // either way, and a throttled table should not take the rest. A put - // already sent runs on ctx, not gctx, so a sibling's failure never cuts - // it off: it answers before the undo below. The caller's cancellation or - // the put's own deadline can, and DynamoDB may then apply it after its - // release. + // The first failure, or the caller's cancellation, skips the puts not + // yet sent: the Reserve fails either way, and a throttled table should + // not take the rest. A put already sent runs on sendCtx, which neither a + // sibling's failure nor the caller's cancellation reaches, so it answers + // before the undo below. Only its own Timeout (call) can cut it off, and + // DynamoDB may then apply it after its release. + sendCtx := context.WithoutCancel(ctx) g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { @@ -354,7 +353,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati return err } sent[i] = true - status, err := s.reserve(ctx, k, token, nowSec, exp) + status, err := s.reserve(sendCtx, k, token, nowSec, exp) if err != nil { return err } @@ -365,7 +364,13 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati return nil }) } - if err := g.Wait(); err != nil { + err := g.Wait() + if err == nil { + // Every sent put answered, but a caller that went away mid-Reserve + // will never commit or release what it claimed. + err = ctx.Err() + } + if err != nil { // A sent put that errored may have landed anyway; releasing a key // its token does not hold is a no-op. Best effort: a put applied // after this, or a release that fails, lapses with the lease. @@ -375,7 +380,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati undo = append(undo, tried[i]) } } - _ = s.Release(context.WithoutCancel(ctx), undo) + _ = s.Release(sendCtx, undo) return nil, err } return claims, nil diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 03a7196f..99ac8b9a 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -633,34 +633,76 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { assert.Equal(t, 2*batchWriteMax, written, "a failed chunk does not cancel the others: their records are published") } -// A put cancelled by its caller is not an answer from the table: it does -// not reset the breaker's count of throttled puts. -func TestDynamo_CallerCancelDoesNotResetBreaker(t *testing.T) { +// A caller that cancels mid-Reserve (a client disconnecting) leaves nothing +// claimed. The fake applies a put after a delay whatever the caller does, as +// DynamoDB applies a request already on the wire, so a put abandoned on the +// cancel would land after its release and hold its id for the lease. +func TestDynamo_CallerCancelLeavesNothingClaimed(t *testing.T) { t.Parallel() - var hang atomic.Bool - started := make(chan struct{}, 1) - _, m := openFake(t, &fakeDynamo{put: func(ctx context.Context, _ *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { - if hang.Load() { + t.Run("a put still unsent", func(t *testing.T) { + t.Parallel() + assertCancelLeavesNothing(t, keys("k0", "k1", "k2")) + }) + t.Run("every put sent", func(t *testing.T) { + t.Parallel() + assertCancelLeavesNothing(t, keys("k0", "k1")) + }) +} + +// assertCancelLeavesNothing cancels a Reserve of ks once two puts are sent. +func assertCancelLeavesNothing(t *testing.T, ks []Key) { + t.Helper() + var ( + mu sync.Mutex + table = map[string]string{} + sent []string + applies sync.WaitGroup + ) + started := make(chan struct{}, 2) + _, m := openFakeWith(t, &fakeDynamo{ + put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := idOf(in.Item[attrKey]) + tk := string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) + mu.Lock() + sent = append(sent, id) + mu.Unlock() + applied := make(chan struct{}) + applies.Go(func() { + time.Sleep(20 * time.Millisecond) + mu.Lock() + table[id] = tk + mu.Unlock() + close(applied) + }) started <- struct{}{} - <-ctx.Done() - return nil, ctx.Err() - } - return nil, &types.ProvisionedThroughputExceededException{} - }}) - for range breakerTrips - 1 { - _, err := m.Reserve(t.Context(), keys("a"), time.Minute) - require.ErrorIs(t, err, ErrUnavailable) - } - hang.Store(true) + select { + case <-applied: + return &dynamodb.PutItemOutput{}, nil + case <-ctx.Done(): + return nil, ctx.Err() + } + }, + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := idOf(in.Key[attrKey]) + tk := string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value) + mu.Lock() + defer mu.Unlock() + if table[id] != tk { + return nil, &types.ConditionalCheckFailedException{} + } + delete(table, id) + return &dynamodb.DeleteItemOutput{}, nil + }, + }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) ctx, cancel := context.WithCancel(t.Context()) - go func() { <-started; cancel() }() - _, err := m.Reserve(ctx, keys("a"), time.Minute) + go func() { <-started; <-started; cancel() }() + _, err := m.Reserve(ctx, ks, time.Minute) require.ErrorIs(t, err, context.Canceled) - hang.Store(false) - _, err = m.Reserve(t.Context(), keys("a"), time.Minute) - require.ErrorIs(t, err, ErrUnavailable) - _, err = m.Reserve(t.Context(), keys("a"), time.Minute) - require.ErrorIs(t, err, errBreakerOpen, "the cancelled put did not reset the count") + applies.Wait() + mu.Lock() + defer mu.Unlock() + assert.Len(t, sent, 2, "a put not yet sent is skipped") + assert.Empty(t, table, "every applied put was released") } func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { From 07e4f77d30dd7200a06a9999d0fffb266a405fb7 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:31:07 -0400 Subject: [PATCH 101/108] fix(dedupe): name ErrUnavailable "dedupe store unavailable" A throttle wraps it too, and logged as "dedupe store is not open". The classify comment now says what ingest answers today: 500 for both, until #629 maps ErrUnavailable to a 503. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- internal/dedupe/dynamodb.go | 5 +++-- internal/dedupe/managed.go | 2 +- tests/integration/dedupe_dynamodb_test.go | 4 ++-- 3 files changed, 6 insertions(+), 5 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index e9e8cd47..88dd557b 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -574,8 +574,9 @@ func newToken() string { // classify maps a DynamoDB error onto the contract: a condition failure is // returned as is for the caller to read, anything retrying later can cure -// wraps ErrUnavailable (503), and the rest — a missing table, denied access, -// a malformed request — is a configuration bug (500). +// wraps ErrUnavailable, and the rest — a missing table, denied access, a +// malformed request — is a configuration bug. Ingest answers both 500 until +// #629 maps ErrUnavailable to a retryable 503. func classify(op string, err error) error { if err == nil { return nil diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index cd42aeb3..09712c50 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -23,7 +23,7 @@ var ErrDisabled = errors.New("dedupe is disabled") // but the store failed to open, and wrapped by a backend's error when a // retry later can succeed. Ingest fails closed on it — the settings asked // for dedupe, so publishing un-deduped is not a fallback. -var ErrUnavailable = errors.New("dedupe store is not open") +var ErrUnavailable = errors.New("dedupe store unavailable") // hashedIDCounter counts ids stored as their SHA-256 (Key.Hashed): an id // longer than MaxIDBytes is a producer sending something other than an id. diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go index 69be6b16..a3cf576b 100644 --- a/tests/integration/dedupe_dynamodb_test.go +++ b/tests/integration/dedupe_dynamodb_test.go @@ -194,7 +194,7 @@ func TestDedupeDynamo_Throttled(t *testing.T) { m := d.Tenant("acme") require.NoError(t, m.Apply(true)) _, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "e1"}}, time.Minute) - require.ErrorIs(t, err, dedupe.ErrUnavailable, "a throttle is worth retrying: 503") + require.ErrorIs(t, err, dedupe.ErrUnavailable, "a throttle is worth retrying") assert.Equal(t, int64(2), sent.Load(), "the SDK retried it once first") }) } @@ -257,7 +257,7 @@ func TestDedupeDynamo_ConfigErrorsAreNotUnavailable(t *testing.T) { require.Error(t, err) var missing *types.ResourceNotFoundException assert.ErrorAs(t, err, &missing) - assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a missing table is a config bug: 500, not 503") + assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a missing table is a config bug, not worth retrying") require.Error(t, d.Check(t.Context())) } From 138b166186718d74da9331fd0d0523dd8c059830 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:36:14 -0400 Subject: [PATCH 102/108] feat(app): refuse boot only for a misconfigured DynamoDB table in use Over a flat settings directory, boot was refused on any failed table check. Now it is refused only when the failure is a misconfiguration (not ErrUnavailable: a missing table, the wrong key schema, access denied) and a tenant has dedupe on. A transient failure, or a misconfigured table no tenant uses yet, boots with the switched-on stores closed and the check retried in the background, as a nested directory already did; a tenant a reload switches on fails closed until it passes. The misconfiguration is logged at ERROR. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- internal/app/dedupe_dynamodb_test.go | 123 ++++++++++++++++++++---- internal/app/wire_dynamodb.go | 38 ++++++-- 6 files changed, 134 insertions(+), 35 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5d405dc6..098224c6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 579f5a21..223b3cec 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -95,7 +95,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. - **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. `wireDedupe`'s `dynamodb` case is `wireDynamoDedupe`, in wire_dynamodb.go (below). `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. -- **wire_dynamodb.go** — `wireDedupe`'s `dynamodb` case, split out of wire.go so the e2e suite's coverage exclude for it (the e2e binary always runs Pebble dedupe, never DynamoDB) doesn't have to blanket wire.go itself: builds the same `dedupe.Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. +- **wire_dynamodb.go** — `wireDedupe`'s `dynamodb` case, split out of wire.go so the e2e suite's coverage exclude for it (the e2e binary always runs Pebble dedupe, never DynamoDB) doesn't have to blanket wire.go itself: builds the same `dedupe.Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on. Boot is refused only for a misconfigured table (an error that is not `ErrUnavailable`) over a flat directory whose tenant has dedupe on; every other failure boots with the switched-on stores closed, the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher, and a flat one's table can come good with no settings change). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index bd1781d0..4c23d2d9 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -61,7 +61,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A tenant a reload switches dedupe on for fails closed the same way until then. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index ddf37795..1aea0f34 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or refuses the pod's credentials refuses boot over a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a throttle or network failure, a nested directory, or no tenant with dedupe on — the pod boots, every tenant with dedupe on (now or after a reload) fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index 310d1140..b2ab5ab5 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -1,8 +1,10 @@ package app import ( + "bytes" "context" "io" + "log/slog" "net" "net/http" "net/http/httptest" @@ -23,14 +25,21 @@ import ( // fakeDynamo answers the DynamoDB JSON protocol for one table, enough for // boot's check, the dev create path, and a claim and its commit. Whether the -// table exists, and whether the endpoint hangs (every call, or one op -// alone), are the test's to switch. +// table exists, whether every call is throttled, and whether the endpoint +// hangs (every call, or one op alone), are the test's to switch. type fakeDynamo struct { - mu sync.Mutex - exists bool - hangs bool - hangOn string // hang calls of this op alone, once set; "" hangs none this way - calls []string + mu sync.Mutex + exists bool + throttles bool + hangs bool + hangOn string // hang calls of this op alone, once set; "" hangs none this way + calls []string +} + +func (f *fakeDynamo) setThrottles(v bool) { + f.mu.Lock() + defer f.mu.Unlock() + f.throttles = v } func (f *fakeDynamo) setExists(v bool) { @@ -76,13 +85,18 @@ func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { if op == "CreateTable" { f.exists = true } - exists, hang := f.exists, f.hangs || op == f.hangOn + exists, throttled, hang := f.exists, f.throttles, f.hangs || op == f.hangOn f.mu.Unlock() if hang { <-r.Context().Done() return } w.Header().Set("Content-Type", "application/x-amz-json-1.0") + if throttled { + w.WriteHeader(http.StatusBadRequest) + _, _ = io.WriteString(w, `{"__type":"com.amazonaws.dynamodb.v20120810#ThrottlingException","message":"Rate exceeded"}`) + return + } if !exists { w.WriteHeader(http.StatusBadRequest) _, _ = io.WriteString(w, `{"__type":"com.amazonaws.dynamodb.v20120810#ResourceNotFoundException","message":"Requested resource not found"}`) @@ -154,20 +168,35 @@ func TestNew_DynamoDBDedupeCreatesTheTableOnlyWhenAsked(t *testing.T) { assert.True(t, a.dedup.For(tenant.Default).Open()) } -// A table that fails the check follows the registry's rule for the shape, -// as a Pebble instance that cannot open does. +// A misconfigured table refuses boot only over a flat directory in which a +// tenant has dedupe on; every other failure boots and fails closed. func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { - t.Run("flat refuses boot", func(t *testing.T) { - for name, patch := range map[string]map[string]any{"dedupe on": dedupeOn, "dedupe off": nil} { - t.Run(name, func(t *testing.T) { - guardGlobals(t) - cfg := testConfig(t, writeSettings(t, patch)) - dynamoConfig(t, cfg, false) - _, err := New(t.Context(), Options{Config: cfg}) - require.ErrorContains(t, err, "dedupe open") - require.ErrorContains(t, err, "ResourceNotFoundException") - }) - } + t.Run("flat with dedupe on refuses boot", func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, dedupeOn)) + dynamoConfig(t, cfg, false) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "dedupe open") + require.ErrorContains(t, err, "ResourceNotFoundException") + require.NotErrorIs(t, err, dedupe.ErrUnavailable) + }) + t.Run("flat with dedupe off boots, and fails closed once it is on", func(t *testing.T) { + dir := writeSettings(t, nil) + cfg := testConfig(t, dir) + dynamoConfig(t, cfg, false) + logs := bootLogged(t) + a, err := New(t.Context(), Options{Config: cfg}) + require.NoError(t, err) + t.Cleanup(func() { assert.NoError(t, a.Close(context.Background())) }) + assert.Contains(t, logs.String(), `level=ERROR msg="dedupe: dynamodb table is misconfigured`) + + store := a.dedup.For(tenant.Default) + _, err = dedupetest.Mark(t.Context(), store, eventKey) + require.ErrorIs(t, err, dedupe.ErrDisabled) + rewriteSettings(t, dir, dedupeOn) + a.tenants.Reload("test") + _, err = dedupetest.Mark(t.Context(), store, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on by a reload while the table is missing: closed, not un-deduped") }) t.Run("nested fails closed", func(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil}) @@ -184,6 +213,58 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { }) } +// A transient failure (a throttle) never refuses boot, even over a flat +// directory with dedupe on: the tenant fails closed until the background +// retry's check passes. +func TestRun_DynamoDBDedupeFlatThrottledRecovers(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, true) + fake.setThrottles(true) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + store := a.dedup.For(tenant.Default) + require.False(t, store.Open()) + _, err = dedupetest.Mark(t.Context(), store, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on, table throttled: ingest fails closed") + + _, stop := runApp(t, a, ln) + fake.setThrottles(false) + require.Eventually(t, store.Open, 10*time.Second, 50*time.Millisecond, "the retry opened the store") + _, err = dedupetest.Mark(context.Background(), store, eventKey) + require.NoError(t, err) + require.NoError(t, stop()) +} + +// bootLogged sends the default logger to a buffer for the rest of the test, +// for a boot that logs what it tolerated. +func bootLogged(t *testing.T) *lockedBuffer { + t.Helper() + guardGlobals(t) + buf := &lockedBuffer{} + slog.SetDefault(slog.New(slog.NewTextHandler(buf, nil))) + return buf +} + +// lockedBuffer is a bytes.Buffer safe for the background retry's logging. +type lockedBuffer struct { + mu sync.Mutex + buf bytes.Buffer +} + +func (b *lockedBuffer) Write(p []byte) (int, error) { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.Write(p) +} + +func (b *lockedBuffer) String() string { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.String() +} + // The reload hook runs under the lock that serializes reloads, so it never // calls DynamoDB: against a table that hangs, a reload returns at once, and a // tenant it switches on fails closed rather than publishing un-deduped. diff --git a/internal/app/wire_dynamodb.go b/internal/app/wire_dynamodb.go index 420111b5..68f68ce4 100644 --- a/internal/app/wire_dynamodb.go +++ b/internal/app/wire_dynamodb.go @@ -19,13 +19,15 @@ var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") // every tenant and every process shares (dedupe.Dynamo), so a tenant's store // opens for free once the table has passed its check. Boot checks it (after // creating it, with create_table on dynamodb-local) whether or not any tenant -// has dedupe on, and never creates it otherwise. A table that fails the check -// follows the registry's rule for the shape, as Pebble's instance does: a -// flat directory refuses boot; a nested one boots with every switched-on -// store closed, so its ingest fails closed. Unlike a local disk, a remote -// table's failure is usually brief (a throttle, credentials not yet issued -// mid-rollout), and a nested directory has no watcher to reload it, so the -// check is then retried in the background, with backoff, until it passes. +// has dedupe on, and never creates it otherwise. Boot is refused only when the +// table is misconfigured (a failure that is not ErrUnavailable: missing, the +// wrong key schema, access denied) over a flat directory in which a tenant +// has dedupe on. Otherwise — a transient failure, a nested directory, or no +// tenant deduping yet — the process boots with every switched-on store +// closed, so its ingest fails closed, and the check is retried in the +// background, with backoff, until it passes: a remote table's failure is +// often brief, a nested directory has no watcher to reload it, and a fixed +// table is picked up without a restart. // The check is network I/O, so the AfterAdopt hook never runs it: the hook // holds the lock that serializes reloads. It applies every store against the // last check's result and wakes the retry, so a reload still retries at once. @@ -94,11 +96,17 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { } }) if err := check(ctx); err != nil { - if !a.tenants.Nested() { + misconfigured := !errors.Is(err, dedupe.ErrUnavailable) + if misconfigured && !a.tenants.Nested() && a.anyDedupeEnabled() { return fmt.Errorf("dedupe open: %w", err) } - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", - "table", c.Table, "error", err) + if misconfigured { + slog.Error("dedupe: dynamodb table is misconfigured; ingest with dedupe on fails closed until it is fixed", + "table", c.Table, "error", err) + } else { + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", + "table", c.Table, "error", err) + } a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { select { @@ -123,3 +131,13 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { apply() return nil } + +// anyDedupeEnabled reports whether a served tenant has dedupe switched on. +func (a *App) anyDedupeEnabled() bool { + for _, store := range a.tenants.All() { + if store.DedupeEnabled() { + return true + } + } + return false +} From 23ac57a3bdd3b88aef1dc5f257ea669b1b1c6313 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:43:40 -0400 Subject: [PATCH 103/108] test(dedupe): order the cancel test's puts after the cancel, not a sleep No put applies before the caller cancels, so neither subtest depends on scheduling. The CHANGELOG line now names what a cancel can still leave held: a put cut off by its own timeout, or a failed release. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- internal/dedupe/dynamodb_test.go | 9 ++++++--- 2 files changed, 7 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1f5e361..0c920053 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` leaves nothing claimed: puts already sent run to their answer and are then released, so its retry is not answered `InFlight` for the lease. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` no longer strands its claim: puts already sent run to their answer on a context its cancellation does not reach, and are then released, so its retry is not answered `InFlight` for the lease. Only a put cut off by its own call timeout (which DynamoDB may apply after the release), or a release that fails, still holds its id until the lease ends. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 99ac8b9a..a6113259 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -636,7 +636,8 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { // A caller that cancels mid-Reserve (a client disconnecting) leaves nothing // claimed. The fake applies a put after a delay whatever the caller does, as // DynamoDB applies a request already on the wire, so a put abandoned on the -// cancel would land after its release and hold its id for the lease. +// cancel would land after its release and hold its id for the lease. No +// put applies before the cancel, so neither case depends on timing. func TestDynamo_CallerCancelLeavesNothingClaimed(t *testing.T) { t.Parallel() t.Run("a put still unsent", func(t *testing.T) { @@ -659,6 +660,7 @@ func assertCancelLeavesNothing(t *testing.T, ks []Key) { applies sync.WaitGroup ) started := make(chan struct{}, 2) + cancelled := make(chan struct{}) _, m := openFakeWith(t, &fakeDynamo{ put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := idOf(in.Item[attrKey]) @@ -668,7 +670,8 @@ func assertCancelLeavesNothing(t *testing.T, ks []Key) { mu.Unlock() applied := make(chan struct{}) applies.Go(func() { - time.Sleep(20 * time.Millisecond) + <-cancelled + time.Sleep(5 * time.Millisecond) // lands after the caller left mu.Lock() table[id] = tk mu.Unlock() @@ -695,7 +698,7 @@ func assertCancelLeavesNothing(t *testing.T, ks []Key) { }, }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) ctx, cancel := context.WithCancel(t.Context()) - go func() { <-started; <-started; cancel() }() + go func() { <-started; <-started; cancel(); close(cancelled) }() _, err := m.Reserve(ctx, ks, time.Minute) require.ErrorIs(t, err, context.Canceled) applies.Wait() From 8ac2fca949789055d9b389de19f05f523c6e5bff Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:04:15 -0400 Subject: [PATCH 104/108] fix(dedupe): classify create_table errors; scope the boot rule in docs CreateTable returned its errors unclassified, so with create_table on a dynamodb-local not listening yet read as a misconfiguration and refused boot over a flat directory. Its errors now go through classify, and a test pins that a transient create failure boots and is retried. settings-directory.mdx still said any failed open refuses boot; that is now scoped to Pebble, with the DynamoDB rule stated beside it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/dedupe_dynamodb_test.go | 21 ++++++++++++++++++++ internal/dedupe/dynamodb.go | 9 +++++---- 4 files changed, 28 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 3adb6298..7fb2552b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot (an endpoint not up yet is a transient failure, retried like the check) and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` does not strand its claim: puts already sent run to their answer on a context its cancellation does not reach, and are then released, so its retry is not answered `InFlight` for the lease. Only a put cut off by its own call timeout (which DynamoDB may apply after the release), or a release that fails, still holds its id until the lease ends. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index e1f0be88..f69d2bbc 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until it opens — the files asked for dedupe, so publishing un-deduped is not a fallback. With `dedupe.backend: pebble` that is the next reload or restart, and at boot a failed open over a flat directory refuses to start, like every other store; with `dynamodb` it is the background retry described below. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part, in either shape: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes. Only a misconfigured table (missing, the wrong key schema, access denied) over a flat directory whose tenant has dedupe on refuses boot instead ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. While its record is being published, an id is held for its lease ([`dedupe.lease`](/configuration#dedupe), 30 seconds by default): another request carrying the same id meanwhile gets `503` (`a request with the same dedupe id is in flight`) with the lease, in whole seconds, as `Retry-After` — see [the ingest errors](/api#post-v1ingesttabletable--ingest-data). An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease, and a later retry of it is accepted again — counted by `wavehouse_ingest_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index b2ab5ab5..e9dd5290 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -237,6 +237,27 @@ func TestRun_DynamoDBDedupeFlatThrottledRecovers(t *testing.T) { require.NoError(t, stop()) } +// With create_table on, an endpoint that fails transiently (dynamodb-local +// still starting) boots too, and the retry creates the table once it answers. +func TestRun_DynamoDBDedupeFlatCreateTableRetries(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, false) + cfg.Dedupe.DynamoDB.CreateTable = true + fake.setThrottles(true) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + store := a.dedup.For(tenant.Default) + require.False(t, store.Open()) + + _, stop := runApp(t, a, ln) + fake.setThrottles(false) + require.Eventually(t, store.Open, 10*time.Second, 50*time.Millisecond, "the retry created the table and opened the store") + assert.True(t, fake.called("CreateTable")) + require.NoError(t, stop()) +} + // bootLogged sends the default logger to a buffer for the rest of the test, // for a boot that logs what it tolerated. func bootLogged(t *testing.T) *lockedBuffer { diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 8117734a..18f21232 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -267,7 +267,8 @@ func (d *Dynamo) Check(ctx context.Context) error { var ErrCreateTableNeedsEndpoint = errors.New("dedupe: create_table is for dynamodb-local only; set the endpoint") // CreateTable creates the table on dynamodb-local, with TTL on ex, and waits -// for it. A table that already exists is left as it is. +// for it. A table that already exists is left as it is. Its errors are +// classified as every call's are, so an endpoint not up yet is ErrUnavailable. func (d *Dynamo) CreateTable(ctx context.Context) error { if d.cfg.Endpoint == "" { return ErrCreateTableNeedsEndpoint @@ -283,17 +284,17 @@ func (d *Dynamo) CreateTable(ctx context.Context) error { return nil } if err != nil { - return fmt.Errorf("dedupe: create table %s: %w", d.cfg.Table, err) + return fmt.Errorf("dedupe: create table %s: %w", d.cfg.Table, classify("create_table", err)) } if err := dynamodb.NewTableExistsWaiter(d.api).Wait(ctx, &dynamodb.DescribeTableInput{TableName: &d.cfg.Table}, time.Minute); err != nil { - return fmt.Errorf("dedupe: wait for table %s: %w", d.cfg.Table, err) + return fmt.Errorf("dedupe: wait for table %s: %w", d.cfg.Table, classify("describe_table", err)) } _, err = d.api.UpdateTimeToLive(ctx, &dynamodb.UpdateTimeToLiveInput{ TableName: &d.cfg.Table, TimeToLiveSpecification: &types.TimeToLiveSpecification{AttributeName: aws.String(attrExpiry), Enabled: aws.Bool(true)}, }) if err != nil { - return fmt.Errorf("dedupe: enable ttl on %s: %w", d.cfg.Table, err) + return fmt.Errorf("dedupe: enable ttl on %s: %w", d.cfg.Table, classify("update_time_to_live", err)) } return nil } From 6096ddd61bb40e7f2a25bb86834addf8872c40bb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:13:23 -0400 Subject: [PATCH 105/108] docs(deployment): a mismatched key schema follows the misconfiguration rule The table section still said boot refuses a table whose key schema does not match, in every case. Drop a sentence configuration.mdx said twice. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 4c23d2d9..0895d4ea 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -61,7 +61,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A tenant a reload switches dedupe on for fails closed the same way until then. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 1aea0f34..a078d2a2 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -454,7 +454,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table: it refuses one whose key schema does not match, and logs a warning if TTL is off. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table and logs a warning if TTL is off; a key schema that does not match is a misconfigured table, handled as described below. An example in Terraform. Replace the tags with your own conventions: From bb357268e1c1f1b3f0961d982ff92e198bea9b9d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:28:25 -0400 Subject: [PATCH 106/108] fix(dedupe): reconcile the windowed ingest with the configurable lease The DynamoDB backend and windowed ingest each described the other's absence: ingest now answers an unavailable dedupe store 503, one Reserve carries a window of up to 256 ids, and the lease is dedupe.lease rather than a fixed 30 seconds. State the lease/window rule once, as the boot check applies it, and pin config's copy of the embedded duplicate window to mq's. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/configuration.mdx | 4 ++-- docs/src/content/docs/durability.md | 2 +- internal/app/app_test.go | 2 +- internal/config/backends.go | 3 ++- internal/config/window_test.go | 15 +++++++++++++++ internal/dedupe/dynamodb.go | 4 ++-- internal/mq/embedded.go | 11 ++++++----- 9 files changed, 31 insertions(+), 14 deletions(-) create mode 100644 internal/config/window_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index a061cb31..45cfb1cb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot (an endpoint not up yet is a transient failure, retried like the check) and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; ingest sends a window of up to 256 ids per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot (an endpoint not up yet is a transient failure, retried like the check) and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` does not strand its claim: puts already sent run to their answer on a context its cancellation does not reach, and are then released, so its retry is not answered `InFlight` for the lease. Only a put cut off by its own call timeout (which DynamoDB may apply after the release), or a release that fails, still holds its id until the lease ends. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/config.yaml b/config.yaml index fcdf336d..cd458e10 100644 --- a/config.yaml +++ b/config.yaml @@ -58,7 +58,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) lease: 30s # how long a claimed id stays pending; at most 59s with the embedded mq (lease + ceil(lease) + 1s within its 2m duplicate window) - reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend, and DynamoDB's idle connections per host; the fan-out has no effect yet (ingest sends one id per call) + reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend, and DynamoDB's idle connections per host; ingest sends a window of up to 256 ids per call # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod # region: "" # empty = AWS_REGION diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 0895d4ea..cf746ba1 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -57,11 +57,11 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, the lease plus its own ceiling to the next whole second plus one more second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59s`: a client that obeys `Retry-After` after a publish whose outcome it never learned can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match, never fewer than the SDK's own default (10). Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match, never fewer than the SDK's own default (10). Ingest reserves and commits a window of up to 256 ids per call; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest `503` (`dedupe store unavailable`, `Retry-After: 5`) until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 3cda3ce2..0a473a9b 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -64,7 +64,7 @@ With [deduplication](/settings-directory#deduplication) on, a `200` also means t With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path. It reads 1,024 keys at a time without holding up commits, then re-reads the expired ones and deletes those still expired; a commit waits only for that last step, at most 1,024 point reads and one unsynced write, however many deleted keys the read stepped over. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The duplicate window has to cover more than the lease alone: the `503` for an uncertain publish sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and a claim's expiry can itself round up by a further second on some backends — the invariant the queue configuration and its tests pin is `2×lease + 1s ≤ window`, not just `lease ≤ window`. Two minutes against a 30-second lease clears that with room to spare. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its dedupe lease ([`dedupe.lease`](/configuration#dedupe), 30 seconds by default) rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The duplicate window has to cover more than the lease alone: the `503` for an uncertain publish sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and a claim's expiry can itself round up by a further second on some backends — the invariant is `lease + ceil(lease) + 1s ≤ window` (`ceil` rounding up to the whole second, so `2×lease + 1s` for a whole-second lease), not just `lease ≤ window`, and boot refuses a `dedupe.lease` that breaks it with the embedded queue, so at most 59 seconds. Two minutes against the default 30-second lease clears that with room to spare. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. ## Check your storage before you trust it diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 57646a3e..16136762 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -568,7 +568,7 @@ func TestNew_RefusesALayerWithoutABackend(t *testing.T) { // A Pebble instance that cannot open follows the registry's own rule for the // shape: a flat directory refuses boot, like every other store, and a nested // one fails closed for every tenant with dedupe on, since they share the -// instance — their ingest answers 500 until a reload or a restart opens it — +// instance — their ingest answers 503 until a reload or a restart opens it — // while the process, and every tenant with dedupe off, carries on. func TestNew_DedupeOpenFailure(t *testing.T) { dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}} diff --git a/internal/config/backends.go b/internal/config/backends.go index b2dc0dd0..b9644b4b 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -167,7 +167,8 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { } // embeddedDuplicateWindow is the embedded ingest stream's duplicate window, -// counted from the stored publish. +// counted from the stored publish. It mirrors mq.EmbeddedDuplicateWindow, +// which config must not import; window_test.go pins the two. const embeddedDuplicateWindow = 2 * time.Minute // maxEmbeddedLease is the longest dedupe.lease the duplicate window covers — diff --git a/internal/config/window_test.go b/internal/config/window_test.go new file mode 100644 index 00000000..9d2559cd --- /dev/null +++ b/internal/config/window_test.go @@ -0,0 +1,15 @@ +package config + +import ( + "testing" + + "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/stretchr/testify/assert" +) + +// config must not import internal/mq (it would pull NATS into every +// importer of config), so the lease cap mirrors the window; this pins them. +func TestEmbeddedDuplicateWindow_MatchesMQ(t *testing.T) { + t.Parallel() + assert.Equal(t, mq.EmbeddedDuplicateWindow, embeddedDuplicateWindow) +} diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 18f21232..e6cb55d2 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -579,8 +579,8 @@ func newToken() string { // classify maps a DynamoDB error onto the contract: a condition failure is // returned as is for the caller to read, anything retrying later can cure // wraps ErrUnavailable, and the rest — a missing table, denied access, a -// malformed request — is a configuration bug. Ingest answers both 500 until -// #629 maps ErrUnavailable to a retryable 503. +// malformed request — is a configuration bug. Ingest answers ErrUnavailable +// with a retryable 503 and a configuration bug with a 500. func classify(op string, err error) error { if err == nil { return nil diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index c257ea39..f852a4c2 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -328,11 +328,12 @@ func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { } // EmbeddedDuplicateWindow is how long an ingest queue remembers a -// WithIdempotencyKey key. It must be at least 2*lease + 1s: a claim left to -// lapse after an uncertain publish is republished once the lease ends, but -// the in-flight 503 tells a client to retry only after the FULL lease, so an -// obedient client's retry can land up to ~2*lease after the original -// Reserve; the +1s covers a backend (DynamoDB, for one) that rounds a +// WithIdempotencyKey key. It must be at least lease + ceil(lease) + 1s +// (2*lease + 1s for a whole-second lease), which config checks against +// dedupe.lease at boot. A claim left to lapse after an uncertain publish is +// republished once the lease ends, but the in-flight 503 tells a client to +// retry only after the FULL lease, so an obedient client's retry can land up +// to ~2*lease after the original Reserve; the +1s covers a backend (DynamoDB, for one) that rounds a // claim's expiry up by as much. Only a window at least that long guarantees // this queue still drops the retry's second copy. const EmbeddedDuplicateWindow = 2 * time.Minute From c361832ba3dd80e41e85a2b6c5e2e288ff629d70 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:33:43 -0400 Subject: [PATCH 107/108] docs(dedupe): retention reaches DynamoDB through TTL, not the sweep Committed items carry ex once a finite dedupe.retention applies, the Pebble sweep does not run for the DynamoDB backend, and a throttled or unreachable table answers the same 503 as a store that is not open. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/deployment.md | 4 ++-- docs/src/content/docs/durability.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- 4 files changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 76c5e19f..0fd1abb9 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -293,7 +293,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | -| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store is not open (for example, it failed to open on a reload); `Retry-After: 5`. Nothing was published, so the retry is safe | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now: it is not open (for example, it failed to open on a reload), or a DynamoDB table is throttling, timing out or unreachable; `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown, other than a full queue or an unreachable broker (below): the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease ([`dedupe.lease`](/configuration#dedupe), 30 seconds by default) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue's duplicate window (two minutes) covers up to ~2×lease plus a margin, not just the lease itself, so a retry timed off `Retry-After` anywhere in this flow stores no second copy; a much later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index adc6cf7c..c5119c1d 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -456,7 +456,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table and logs a warning if TTL is off; a key schema that does not match is a misconfigured table, handled as described below. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. TTL removes lapsed claims, and a committed id once its [`dedupe.retention`](/settings-directory#deduplication) ends. With the default retention `"0"` (forever) a committed item carries no `ex` and is kept, so the table grows by one item (about 200 bytes) per distinct id. Boot checks the table and logs a warning if TTL is off; a key schema that does not match is a misconfigured table, handled as described below. An example in Terraform. Replace the tags with your own conventions: @@ -520,7 +520,7 @@ For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for exam - **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. -- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. +- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table until its retention ends, forever at the default (see TTL above), at DynamoDB's per-GB-month rate. - **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five throttled or unreachable claims in a row within one second, the backend stops calling the table for a second and fails every tenant's dedupe requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). A duplicate or in-flight answer is not a failure and resets the count. - **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`, `wavehouse_dedupe_dynamodb_short_circuits_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 0a473a9b..a67b3a75 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_ingest_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path. It reads 1,024 keys at a time without holding up commits, then re-reads the expired ones and deletes those still expired; a commit waits only for that last step, at most 1,024 point reads and one unsynced write, however many deleted keys the read stepped over. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. +With a finite `dedupe.retention` and `dedupe.backend: pebble`, expired ids are deleted by a background sweep, an hour apart; on DynamoDB the table's TTL deletes them instead ([Deployment](/deployment#a-shared-dedupe-table-on-dynamodb)). Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path. It reads 1,024 keys at a time without holding up commits, then re-reads the expired ones and deletes those still expired; a commit waits only for that last step, at most 1,024 point reads and one unsynced write, however many deleted keys the read stepped over. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its dedupe lease ([`dedupe.lease`](/configuration#dedupe), 30 seconds by default) rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The duplicate window has to cover more than the lease alone: the `503` for an uncertain publish sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and a claim's expiry can itself round up by a further second on some backends — the invariant is `lease + ceil(lease) + 1s ≤ window` (`ceil` rounding up to the whole second, so `2×lease + 1s` for a whole-second lease), not just `lease ≤ window`, and boot refuses a `dedupe.lease` that breaks it with the embedded queue, so at most 59 seconds. Two minutes against the default 30-second lease clears that with room to spare. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index a06ac1c8..051791d7 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -190,7 +190,7 @@ Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.ba - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until it opens — the files asked for dedupe, so publishing un-deduped is not a fallback. With `dedupe.backend: pebble` that is the next reload or restart, and at boot a failed open over a flat directory refuses to start, like every other store; with `dynamodb` it is the background retry described below. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part, in either shape: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes. Only a misconfigured table (missing, the wrong key schema, access denied) over a flat directory whose tenant has dedupe on refuses boot instead ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. While its record is being published, an id is held for its lease ([`dedupe.lease`](/configuration#dedupe), 30 seconds by default): another request carrying the same id meanwhile gets `503` (`a request with the same dedupe id is in flight`) with the lease, in whole seconds, as `Retry-After` — see [the ingest errors](/api#post-v1ingesttabletable--ingest-data). An id is committed only after its record is published; if that commit fails (counted by `wavehouse_ingest_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.retention` (optional; seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed, and a `config.json` without the key means the same. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.retention` (optional; seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed, and a `config.json` without the key means the same. Once an id's retention has ended, the next record carrying it is published as new. With `dedupe.backend: pebble`, a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. With `dynamodb`, no sweep runs: the table's TTL on `ex` deletes the item ([Deployment](/deployment#a-shared-dedupe-table-on-dynamodb)). A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. - `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest, so a table with no `retention` keeps the tenant's (forever when the tenant sets none). A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse From 64902034e5be69934c4cf602f18f9cc05b8a7241 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 17:29:20 -0400 Subject: [PATCH 108/108] docs: name the shared dedupe backend where the merge left it out The multiple-instances section from #614 said dedupe is always per instance, and the boot-config list named only cache.redis. Both now name dedupe.dynamodb. The integration setup also starts dynamodb-local. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- AGENTS.md | 2 +- docs/src/content/docs/deployment.md | 4 ++-- docs/src/content/docs/development.md | 4 ++-- docs/src/content/docs/settings-directory.mdx | 2 +- 4 files changed, 6 insertions(+), 6 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 32581d83..4573713a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -153,7 +153,7 @@ If `make ci` passes locally, your commit has crossed the same gates CI will run ### Running `make ci` (for agents) -`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse and a Redis via **testcontainers on random host ports**, and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). +`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse and a Redis via **testcontainers on random host ports** (the integration suite also dynamodb-local), and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). Run it via the **background Bash tool** (`run_in_background: true`) and wait for the completion notification; the harness re-invokes you on exit, so polling the log with `tail` only burns context: diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 6c986e61..9c8029a4 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -425,9 +425,9 @@ The folder name is the tenant id, and each folder is a complete settings directo ## Multiple instances and the shared cache -Several WaveHouse instances can serve one ClickHouse behind a load balancer, but most of what each one holds is its own. The message queue is embedded, so an event is inserted by the instance that took its `POST /v1/ingest`, and reaches only that instance's SSE subscribers. The dedupe store is per instance too, so an id one instance has seen is new to another. +Several WaveHouse instances can serve one ClickHouse behind a load balancer, but most of what each one holds is its own. The message queue is embedded, so an event is inserted by the instance that took its `POST /v1/ingest`, and reaches only that instance's SSE subscribers. With the default `dedupe.backend: pebble` the dedupe store is per instance too, so an id one instance has seen is new to another; [`dedupe.backend: dynamodb`](#a-shared-dedupe-table-on-dynamodb) shares seen ids across instances. -The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. The server is a standalone one or a Redis Cluster; Sentinel (`mode: sentinel`) refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656), since the cache does not yet authenticate to the sentinels or refresh their topology. +The query-result cache and the dedupe store are the layers that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. The server is a standalone one or a Redis Cluster; Sentinel (`mode: sentinel`) refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656), since the cache does not yet authenticate to the sentinels or refresh their topology. **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index c9052d62..a3d5b0d3 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -16,7 +16,7 @@ You need these on your `PATH` before any `make` recipe will work end-to-end: | **Go** | 1.26+ (matches `go.mod`) | Compiles `cmd/wavehouse`; also runs the pinned `tool` deps (`gotestsum`, `gofumpt`, `goimports`, `govulncheck`, `deadcode`, `gsa`, `goda`) via `go tool` | [go.dev/dl](https://go.dev/dl/) | | **GNU Make** | **4.0+** | The Makefile uses `--output-sync=target` (Make 4 only) and bash-pinned recipes. macOS ships with BSD Make 3.81, which **will not work** | macOS: `brew install make` then use `gmake` or put `$(brew --prefix make)/libexec/gnubin` on your PATH. Linux: usually already installed | | **bash** | 4+ recommended | Recipes are pinned to `bash`; the helper scripts under `scripts/` use `set -euo pipefail` and bash arrays | macOS default is bash 3.2 (works for current recipes, but `brew install bash` is safer); Linux distros ship 4+ | -| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also runs the shared cache backend against Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | +| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), the integration suite also dynamodb-local, and the integration suite also runs the shared cache backend against Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | | **Node.js** | 22 LTS — pinned via `.nvmrc` at the repo root | Runtime for pnpm and the Vitest suites. Pinned to match CI (`setup-node` uses 22) and to avoid Node-major surprises; older Vitest versions in this repo were known to crash on Node 26 with a V8 heap-allocation abort | [nodejs.org](https://nodejs.org/) or `nvm use` / `fnm use` / `volta` (all read `.nvmrc`) | | **pnpm** | 11.21+ (pinned via `packageManager` in the root `package.json`) | Package manager for the TypeScript SDK, E2E test harness, and docs site (managed as a single pnpm workspace from the repo root); `make build-ts`, `make test-ts`, `make test-e2e`, `make build-docs`, `make dev-docs`, `make preview-docs` all shell out to `pnpm` | `corepack enable && corepack prepare pnpm@11.21.0 --activate` (recommended), or `npm i -g pnpm` | | **git** + **curl** | any recent | `git` for source + version metadata in builds; `curl` is used by the Makefile to fetch the pinned `golangci-lint` binary into `.bin/` | usually preinstalled | @@ -345,7 +345,7 @@ Each test target writes `covdata` to `tmp/coverage//data/`, renders a tex | E2E tests (SDK) | `tests/e2e/sdk/*.test.ts` | Yes | `make test-e2e` | - **Unit tests** live beside the code they test (e.g., `internal/discovery/discovery_test.go`). They use mocks or embedded NATS (in-process, no Docker needed). -- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. `shared_cache_test.go` starts its own Redis testcontainer per test (`startRedis`) and boots extra, independent `cache.backend: redis` instances over that same ClickHouse (`bootRedisApp`), to exercise the cache shared across processes rather than one package in isolation. +- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and a dynamodb-local one (for the DynamoDB dedupe backend's tests), and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. `shared_cache_test.go` starts its own Redis testcontainer per test (`startRedis`) and boots extra, independent `cache.backend: redis` instances over that same ClickHouse (`bootRedisApp`), to exercise the cache shared across processes rather than one package in isolation. Shared test utilities live in `internal/testutil/`. The packages log through `slog.Default()`, so tests reach log output through `internal/testutil/logtest`: `logtest.Silence()` in a package's `TestMain` discards it, and `logtest.Capture(t, level)` routes it to a buffer for a test that asserts on log lines — such a test must not call `t.Parallel()`, because the default logger is process-wide. A test that starts the embedded broker keeps its store in `internal/testutil/storedir`'s `storedir.New(t)` rather than a bare `t.TempDir()` (`testutil.NewEmbeddedMQ` does): the NATS server can finish writing a consumer's state after `Close` returns, which fails `t.TempDir`'s one-shot removal, and `storedir` removes the store again until those writes have landed ([#442](https://github.com/Wave-RF/WaveHouse/issues/442)). diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c16e9dbe..0a390de4 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -181,7 +181,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`) and a shared backend's connection (`cache.redis`), the process's `roles`, resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`) and a shared backend's connection (`cache.redis`, `dedupe.dynamodb`), how a dedupe claim behaves (`dedupe.lease`, `dedupe.reserve_concurrency`), the process's `roles`, resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication