From 37ef7578002562f34a0f9f2ce25c707ff80e585f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:11:07 -0400 Subject: [PATCH 01/22] fix(ingest): windowed reserve/publish/commit; 503 when dedupe is unavailable Each window of up to 256 records is prepared, reserved in one dedupe call, published in order and committed in one call. Deduped records are published under an idempotency key (Nats-Msg-Id), and each tenant's ingest stream keeps an explicit two-minute duplicate window, so a publish whose outcome is unknown leaves its claim to lapse and the retry's copy is dropped by the queue. A dedupe store that cannot answer is a 503 with Retry-After: 5. Fixes #384. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 1 + docs/src/content/docs/api.md | 12 +- docs/src/content/docs/architecture.md | 21 +- docs/src/content/docs/durability.md | 6 + docs/src/content/docs/sdk/reference.md | 2 +- docs/src/content/docs/settings-directory.mdx | 4 +- internal/api/ingest.go | 338 +++++++++----- internal/api/ingest_seams.go | 4 +- internal/api/ingest_test.go | 122 ++++-- internal/api/ingest_window_test.go | 435 +++++++++++++++++++ internal/dedupe/key.go | 9 + internal/dedupe/key_test.go | 33 ++ internal/mq/embedded.go | 17 +- internal/mq/embedded_test.go | 29 ++ internal/mq/mq.go | 14 + internal/testutil/mocks.go | 31 +- 17 files changed, 899 insertions(+), 183 deletions(-) create mode 100644 internal/api/ingest_window_test.go create mode 100644 internal/dedupe/key_test.go diff --git a/AGENTS.md b/AGENTS.md index 636156865..ffbd3c16f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` +- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal; `WithIdempotencyKey` makes a republish inside the queue's duplicate window a no-op), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) - **`policy/`** — Hasura-style access control, **role-first**: `TablePolicy` is `map[string]RolePermissions`, and a role's grant splits by operation into `SelectPermissions` (columns, row `filter`, aggregations, the `max_*` limits) and `InsertPermissions` (columns, `check`) — so a field only one side honors does not exist on the other. `Evaluate()` resolves ONE operation and leaves the other side **nil** (`Select *ResolvedSelect` / `Insert *ResolvedInsert`), which every accessor fails closed on — nil is "not resolved", distinct from an empty side, which is "unrestricted" (what the admin return builds). Claim templating (`{{ jwt.claim.path }}`) resolves during that call. Policies come from `Source`, a `func() *Policy` read per call (`settings.Store.Policy` in production, `Static(p)` in tests) @@ -58,7 +58,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. 6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format`, or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. -8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase — `Reserve` → publish → `Commit`, or `Release` when the publish fails; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. +8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase, one call per phase per window of up to 256 records — `Reserve` → publish (under the id's idempotency key) → `Commit`, or `Release` when the publish definitely failed, while one whose outcome is unknown is left to lapse; a store that cannot answer is a `503` + `Retry-After`; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 56aa5a3ae..041ef99e4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,6 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and the retry after the lease is dropped by the queue if the first copy was stored. A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index dc79ed6b9..11f9f1f07 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -272,8 +272,9 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now (not open, or a remote backend throttled or unreachable); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate. | +| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored (it remembers the key for two minutes), so the retry never stores a second copy. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | @@ -385,13 +386,14 @@ A `200` is returned whenever the body was read and the records were processed | 403 | `{"error":"forbidden"}` (empty-role variant: `forbidden: request has no role and no public default_role is configured`) | The resolved role lacks `insert` on the table (checked once, before any record) | | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | -| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest | -| 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. As for `publish failed`, the failing record's id is given back and the records before it keep theirs | -| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | +| 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the records before it keep their ids, so a whole-batch retry reports those as duplicates; the failing record's id is left to lapse as on the single-object path, and the rest of its window's ids are given back | +| 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. The records before the refused one keep their ids, and its id and the rest of its window's are given back | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now; `Retry-After: 5`. Nothing in the window being reserved was published; the windows before it were, and keep their ids | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). Nothing in that record's window was published; the windows before it were | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] -A batch aborted partway (a `503`/`500`, a JSON-array syntax error, or an NDJSON line over the 10 MiB line bound, after some leading records were already published) re-publishes those leading records when the whole batch is retried. A whole-body read failure is **not** one of these: a `413`, or the `400 invalid request body` of an upload cut off in transit, is decided before any record is processed, so nothing is published — safe to retry, once split for a `413`. Enable deduplication if duplicate suppression matters — this is the same at-least-once property the single-object path already has (the SDK retries both on `503`). +A batch aborted partway (a `503`/`500`, a JSON-array syntax error, or an NDJSON line over the 10 MiB line bound, after some leading records were already published) re-publishes those leading records when the whole batch is retried. Records are published in windows of 256, in order: a read error or a dedupe failure drops the open window unpublished, so what an aborted batch published is the windows before it, plus, after a publish failure, the records of its window before the failing one. A whole-body read failure is **not** one of these: a `413`, or the `400 invalid request body` of an upload cut off in transit, is decided before any record is processed, so nothing is published — safe to retry, once split for a `413`. Enable deduplication if duplicate suppression matters — this is the same at-least-once property the single-object path already has (the SDK retries both on `503`). ::: **curl example (JSON array):** diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 740669323..eb660c17d 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -80,7 +80,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. -- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup (the id reserved once the record is encoded, committed after the publish, released if the publish fails; an id another request holds answers `503` with the lease as `Retry-After`), and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). +- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates and encodes each record, and runs the records in windows of up to 256 (`ingestWindow`) through three phases: one dedupe `Reserve` for the window's ids, the publishes in record order (a deduped record under `mq.WithIdempotencyKey`, keyed by `dedupe.IdempotencyKey`), and one `Commit` of the published ids — a window is the unit of a dedupe round trip and of Pebble's commit `fsync`. An id another request holds answers `503` with the lease as `Retry-After`, a store that cannot answer (`dedupe.ErrUnavailable`) `503` with `Retry-After: 5`; a publish that fails at a record commits the ones before it and releases the rest, except that a failure other than `mq.ErrQueueFull` may have stored the event, so that record's claim is left to lapse and the idempotency key drops the retry's copy. Each row goes through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. @@ -147,10 +147,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. -- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. +- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape; `WithIdempotencyKey` marks a publish so that a second one carrying the same key inside the queue's duplicate window is dropped and reported as success. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`) and remembering idempotency keys for `EmbeddedDuplicateWindow` (two minutes, which a dedupe lease must not exceed), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline @@ -228,11 +228,16 @@ Client POST /v1/ingest?table={table} an unparseable value passes through verbatim for ClickHouse's parser to judge) → Optional dedupe: resolve the id (configurable ID field; a row missing it or setting it to null is published un-deduped + logged/counted, or rejected - under require_id); once the record is encoded, reserve (tenant, table, id): - a duplicate is skipped, an id another request holds → 503 + Retry-After - (the 30s lease) - → Publish to NATS JetStream (ingest.{tenant}.{table}) - → Commit the reserved id; on a failed publish, release it instead + under require_id) + → Encode the record; the steps below run per window of up to 256 records + → Reserve the window's (tenant, table, id) keys in one call: a duplicate is + skipped, an id another request holds → 503 + Retry-After (the 30s lease), + a store that cannot answer → 503 + Retry-After: 5 + → Publish each record to NATS JetStream (ingest.{tenant}.{table}), a deduped + one under its idempotency key + → Commit the published ids in one call; on a failed publish, commit the + records before it and release the rest (a publish whose outcome is unknown + keeps its claim until the lease lapses) → 200 OK returned immediately → (If the tenant's NATS stream is full, or not open: 503 + Retry-After header, the id released) diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 8e57d8231..a0e8e36ed 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -58,6 +58,12 @@ The strict guarantee translates well to managed cloud infrastructure — the pre The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) is that a single-threaded benchmark looks fine while a concurrent one is far worse — so always benchmark with multiple writers, and benchmark the guest **and** the host if virtualized. +## Deduplication: one more fsync per window + +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids are committed to the dedupe store, and on the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. + +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes. The retry that follows the lease therefore stores no second copy. That holds only while the lease is shorter than the stream's duplicate window. + ## Check your storage before you trust it Replicate JetStream's exact pattern — a 4 KiB write followed by a flush, in a tight loop — and report the percentiles. The numbers that matter are **p99** and **max**: those are your worst-case publish latency. diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index a7a443876..40ca00484 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -32,7 +32,7 @@ The SDK **never throws** for anything the server returns — all API errors come | 403 | `HTTP_403` | No | Insufficient permissions | | 404 | `HTTP_404` | No | Table, pipe, or tenant not found | | 500 | `HTTP_500` | Yes | Server error (retried per `maxRetries`) | -| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | +| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a dedupe store that cannot answer (`dedupe store unavailable`, `Retry-After: 5`), a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | | 0 | `NETWORK_ERROR` | Yes | Network failure (retried with exponential backoff) | | 0 | `ABORTED` | No | Request canceled via `AbortSignal` | | 0 | `SSE_CONNECT_ERROR` | No | Stream could not be started (e.g. a non-absolute `baseURL`) | diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 905ba816c..a9bdeb82b 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,8 +185,8 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index f40515164..acf8bec88 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -37,6 +37,11 @@ import ( // with the admin query handler — see internal/api/query.go. const maxReportedResults = 10000 +// ingestWindow is how many records a batch prepares before reserving, +// publishing and committing them together: one dedupe call per phase per +// window rather than per record, and at most one window of encoded rows held. +const ingestWindow = 256 + // IngestHandler handles POST /v1/ingest?table={table} type IngestHandler struct { // Registry yields the request tenant's schema registry. @@ -68,6 +73,8 @@ type IngestHandler struct { // tests can pin the cap-overflow path without allocating 16 MiB per run; not // a production tuning knob, hence unexported. Mirrors QueryHandler. maxRequestBytes int64 + // window overrides ingestWindow when > 0, for tests and benchmarks. + window int } func NewIngestHandler(registry RegistrySource, pub mq.Publisher) *IngestHandler { @@ -133,11 +140,14 @@ type recordReject struct { // requestAbort is a whole-request failure: this record and every one that // follows is refused. Both paths stop and return the status; the batch path -// abandons the remaining records rather than silently losing the tail. +// abandons the remaining records rather than silently losing the tail. What +// earlier windows published stays published, and with dedupe on stays +// committed, so a whole-batch retry reports those records as duplicates. // // Most causes are TRANSIENT system conditions, where abandoning the tail is what // makes the batch safe to retry: publish backpressure (503), a publish/marshal -// failure (500), a dedup backend error (500), an id another request holds (503). +// failure (500), a dedupe store that cannot answer (503) or fails (500), an id +// another request holds (503). // // One is not. An insert grant that resolved for the other operation is a 403 and // a caller/config bug — retrying cannot help. It aborts rather than rejecting @@ -322,16 +332,21 @@ func (h *IngestHandler) handleSingle( return } - dup, reject, abort := h.processRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) + rec, abort := h.prepareRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) + if abort == nil && rec.reject == nil { + window := []pendingRecord{rec} + abort = h.ingestWindow(ctx, store, table, scope, window) + rec = window[0] + } if abort != nil { writeAbort(w, abort) return } - if reject != nil { - writeJSONError(w, reject.Status, reject.Message) + if rec.reject != nil { + writeJSONError(w, rec.reject.Status, rec.reject.Message) return } - if dup { + if rec.duplicate { w.Header().Set("Content-Type", "application/json") _ = json.NewEncoder(w).Encode(map[string]bool{"duplicate": true}) return @@ -363,6 +378,24 @@ func (h *IngestHandler) handleBatch( checkGuard *recordReject, ) { result := batchResult{Results: []recordResult{}} + size := h.window + if size <= 0 { + size = ingestWindow + } + window := make([]pendingRecord, 0, min(size, 16)) + // flush runs the window's records through reserve → publish → commit and + // reports them in order; false when it aborted the request. + flush := func() bool { + if abort := h.ingestWindow(ctx, store, table, scope, window); abort != nil { + writeAbort(w, abort) + return false + } + for i := range window { + result.add(&window[i]) + } + window = window[:0] + return true + } for { data, err := rr.Next() @@ -372,8 +405,10 @@ func (h *IngestHandler) handleBatch( if err != nil { if rse, ok := errors.AsType[*recordSyntaxError](err); ok { result.Total++ - result.Failed++ - appendResult(&result, recordResult{Index: result.Total, Error: rse.Error()}) + window = append(window, pendingRecord{index: result.Total, reject: &recordReject{Message: rse.Error()}}) + if len(window) == size && !flush() { + return + } continue } // Unreachable while the body is buffered — a bytes.Reader cannot produce @@ -392,26 +427,22 @@ func (h *IngestHandler) handleBatch( } result.Total++ - idx := result.Total - dup, reject, abort := h.processRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) + rec, abort := h.prepareRecord(ctx, store, table, scope, schema, perms, role, data, now, checkGuard) if abort != nil { // Whole-request failure: surface the status rather than recording a // request-scoped condition as per-record loss (see requestAbort). + // Nothing in the open window has been published. writeAbort(w, abort) return } - if reject != nil { - result.Failed++ - appendResult(&result, recordResult{Index: idx, Error: reject.Message}) - continue - } - if dup { - result.Duplicates++ - appendResult(&result, recordResult{Index: idx, Duplicate: true}) - continue + rec.index = result.Total + window = append(window, rec) + if len(window) == size && !flush() { + return } - result.Succeeded++ - appendResult(&result, recordResult{Index: idx, Ok: true}) + } + if len(window) > 0 && !flush() { + return } slog.InfoContext(ctx, "batch ingested", "table", table, @@ -421,12 +452,24 @@ func (h *IngestHandler) handleBatch( _ = json.NewEncoder(w).Encode(result) } -// appendResult records a per-record outcome up to maxReportedResults. The -// batchResult counts are incremented by the caller and stay authoritative even -// when the Results slice is truncated. -func appendResult(result *batchResult, entry recordResult) { - if len(result.Results) < maxReportedResults { - result.Results = append(result.Results, entry) +// add counts rec's outcome and records it up to maxReportedResults; the counts +// stay authoritative when Results is truncated. Total is counted as records +// are read. +func (r *batchResult) add(rec *pendingRecord) { + entry := recordResult{Index: rec.index} + switch { + case rec.reject != nil: + r.Failed++ + entry.Error = rec.reject.Message + case rec.duplicate: + r.Duplicates++ + entry.Duplicate = true + default: + r.Succeeded++ + entry.Ok = true + } + if len(r.Results) < maxReportedResults { + r.Results = append(r.Results, entry) } } @@ -523,19 +566,30 @@ func (h *IngestHandler) policyCheckGuard( } } -// processRecord runs the per-record pipeline shared by the single-object and -// batch ingest paths: schema validation → column/check permission enforcement -// (with claim-derived auto-injection) → optional dedup → publish. The -// table-level insert grant is checked once by the caller before any record is -// processed, so perms here drives only the per-column and per-row checks (it is -// nil when no policy store is configured). data may be mutated to auto-inject -// check-clause values. +// pendingRecord is one record between prepareRecord and its outcome. +type pendingRecord struct { + index int // 1-based position in a batch + reject *recordReject // non-nil: the record is bad and is not published + payload []byte // the encoded envelope to publish + // key is the record's dedupe identity, nil when it is published + // un-deduped; claim is Reserve's answer for it. + key *dedupe.Key + claim dedupe.Claim + duplicate bool +} + +// prepareRecord runs the per-record half of the pipeline shared by the +// single-object and batch ingest paths: schema validation → column/check +// permission enforcement (with claim-derived auto-injection) → timestamp +// canonicalization → dedupe id resolution → encoding. Reserving, publishing +// and committing happen per window, in ingestWindow. The table-level insert +// grant is checked once by the caller before any record is processed, so perms +// here drives only the per-column and per-row checks (it is nil when no policy +// store is configured). data may be mutated to auto-inject check-clause values. // -// Exactly one of the outcomes is meaningful per call: -// - duplicate true: the record was skipped by dedup (reject/abort nil). -// - reject non-nil: the record is bad; the rest of a batch may still proceed. -// - abort non-nil: a whole-request failure; the caller stops and returns it. -func (h *IngestHandler) processRecord( +// A record the rest of a batch may proceed past comes back with reject set; +// abort non-nil is a whole-request failure the caller stops and returns. +func (h *IngestHandler) prepareRecord( ctx context.Context, store *settings.Store, table, scope string, @@ -545,10 +599,10 @@ func (h *IngestHandler) processRecord( data map[string]any, now time.Time, checkGuard *recordReject, -) (duplicate bool, reject *recordReject, abort *requestAbort) { +) (rec pendingRecord, abort *requestAbort) { if err := h.validator().Validate(schema, data); err != nil { slog.WarnContext(ctx, "schema validation failed", "error", err, "table", table) - return false, &recordReject{Status: http.StatusBadRequest, Message: err.Error()}, nil + return pendingRecord{reject: &recordReject{Status: http.StatusBadRequest, Message: err.Error()}}, nil } // DEEP AUTH: column-level allow/deny + check clauses. @@ -556,10 +610,10 @@ func (h *IngestHandler) processRecord( for col := range data { if !perms.IsColumnAllowed(col, true) { slog.WarnContext(ctx, "column insertion forbidden", "column", col, "role", role) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusForbidden, Message: fmt.Sprintf("column %q not allowed for insert", col), - }, nil + }}, nil } } // Through the accessor, not a bare read. The check loop iterates a side's @@ -580,7 +634,7 @@ func (h *IngestHandler) processRecord( // permission failures for one mis-wired grant. slog.ErrorContext(ctx, "insert checks consulted on a grant resolved for another operation", "table", table, "role", role) - return false, nil, &requestAbort{ + return pendingRecord{}, &requestAbort{ Status: http.StatusForbidden, Message: "insert permissions were not resolved for this request", } @@ -595,7 +649,7 @@ func (h *IngestHandler) processRecord( // a record that supplies the column fails schema validation first with // a different message, and a batch should report each its own cause. if checkGuard != nil { - return false, checkGuard, nil + return pendingRecord{reject: checkGuard}, nil } // A []any value is an _in check: the inserted value must be present and // one of the allowed set. Unlike the scalar _eq case there is no single @@ -604,10 +658,10 @@ func (h *IngestHandler) processRecord( actual, ok := data[col] if !ok || !h.checker().InSet(actual, set) { slog.WarnContext(ctx, "check clause failed", "column", col, "allowed", set, "actual", actual, "present", ok) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusForbidden, Message: fmt.Sprintf("check failed for column %q", col), - }, nil + }}, nil } continue } @@ -624,10 +678,10 @@ func (h *IngestHandler) processRecord( // reading the token's own JSON type didn't give it. if !h.checker().Matches(actual, requiredVal) { slog.WarnContext(ctx, "check clause failed", "column", col, "expected", requiredVal, "actual", actual) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusForbidden, Message: fmt.Sprintf("check failed for column %q", col), - }, nil + }}, nil } } else { // Auto-inject the required value if not provided — as a plain @@ -652,23 +706,22 @@ func (h *IngestHandler) processRecord( // always states them, so no compiled fallback is needed), so a reload // lands at a record boundary. A Deduplicator without a settings source is // a wiring bug, not a mode — main wires both or neither. The id is claimed - // only once the record is encoded, so nothing but the publish can fail - // while the claim is held. - var dedupKey *dedupe.Key + // in ingestWindow, once every record of the window is encoded, so nothing + // but the publish can fail while the claim is held. if h.Dedup != nil && h.DedupeSettings != nil { if enabled, idField, requireID := h.DedupeSettings(store, table); enabled { // An explicit null is as missing as an absent key (#370): fmt.Sprint // would make every null "", one id for every such record. if idVal, ok := data[idField]; ok && idVal != nil { - dedupKey = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} + rec.key = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} } else { dedupeMissingIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) if requireID { slog.WarnContext(ctx, "dedupe id_field missing or null; rejecting", "id_field", idField, "table", table) - return false, &recordReject{ + return pendingRecord{reject: &recordReject{ Status: http.StatusBadRequest, Message: fmt.Sprintf("missing dedupe id field %q", idField), - }, nil + }}, nil } slog.WarnContext(ctx, "dedupe id_field missing or null; publishing without idempotency", "id_field", idField, "table", table) } @@ -683,7 +736,7 @@ func (h *IngestHandler) processRecord( row, err := ingest.EncodeCompactRow(cols, data) if err != nil { slog.ErrorContext(ctx, "failed to encode compact row", "error", err, "table", table) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} + return pendingRecord{}, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} } evt := ingest.EventMessage{ @@ -695,104 +748,173 @@ func (h *IngestHandler) processRecord( Row: row, } - payload, err := json.Marshal(evt) + rec.payload, err = json.Marshal(evt) if err != nil { slog.ErrorContext(ctx, "failed to marshal event message", "error", err) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} + return pendingRecord{}, &requestAbort{Status: http.StatusInternalServerError, Message: "marshal failed"} } + return rec, nil +} +// ingestWindow reserves, publishes and commits one window of prepared +// records, in three phases: one Reserve for every keyed record, the publishes +// in record order, one Commit for every claim published. Rejected and +// duplicate records are skipped. It sets each record's outcome and returns an +// abort when the request must stop; what the window published before a failure +// is committed first, so the retry reports it as duplicates (see publishFailed). +func (h *IngestHandler) ingestWindow(ctx context.Context, store *settings.Store, table, scope string, recs []pendingRecord) *requestAbort { var dd dedupe.Deduplicator - var claims []dedupe.Claim - if dedupKey != nil { + var keyed []int + for i := range recs { + if recs[i].reject == nil && recs[i].key != nil { + keyed = append(keyed, i) + } + } + if len(keyed) > 0 { dd = h.Dedup(store) - var duplicate bool - var abort *requestAbort - claims, duplicate, abort = h.reserve(ctx, dd, *dedupKey) - if duplicate || abort != nil { - return duplicate, nil, abort + if abort := h.reserve(ctx, dd, table, recs, keyed); abort != nil { + return abort } } - slog.DebugContext(ctx, "publishing event to the ingest queue", "table", table, "scope", scope) - if err := h.Publisher.Publish(ctx, mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope}, payload); err != nil { - // The record is not in the queue, so its id goes back: the client's - // retry must not read as a duplicate of it (#384). - releaseClaims(ctx, dd, claims) - if errors.Is(err, mq.ErrQueueFull) { - slog.WarnContext(ctx, "ingest queue is full", "tenant", store.Tenant(), "error", err, "table", table, "scope", scope) - return false, nil, &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} + topic := mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope} + for i := range recs { + rec := &recs[i] + if rec.reject != nil || rec.duplicate { + continue + } + var opts []mq.PublishOpt + if rec.claim.Status == dedupe.Claimed { + // The retry of an uncertain publish carries the same id, so the + // queue drops its copy if the first one landed. + opts = append(opts, mq.WithIdempotencyKey(dedupe.IdempotencyKey(store.Tenant(), rec.claim.Key))) + } + if err := h.Publisher.Publish(ctx, topic, rec.payload, opts...); err != nil { + return h.publishFailed(ctx, dd, topic, recs, i, err) } - slog.ErrorContext(ctx, "failed to publish to the ingest queue", "tenant", store.Tenant(), "error", err, "table", table, "scope", scope) - return false, nil, &requestAbort{Status: http.StatusInternalServerError, Message: "publish failed"} } - commitClaims(ctx, dd, claims, table) - return false, nil, nil + commitClaims(ctx, dd, claimedIn(recs), table) + return nil } -// reserve claims key for one record. A duplicate skips the record; a key -// another request holds aborts with 503 and the lease as Retry-After, since -// that request's outcome decides this one's. ErrDisabled — a reload switched -// the store off after the settings snapshot was read — publishes un-deduped, -// as a record under the other setting would have been. -func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, key dedupe.Key) (claims []dedupe.Claim, duplicate bool, abort *requestAbort) { +// reserve claims the keys of recs[keyed] in one call and records each answer. +// A duplicate is skipped. A key another request holds releases the window's +// claims and aborts with 503 and the lease as Retry-After, since that +// request's outcome decides this one's. A store that cannot answer now is a +// 503 too; nothing in the window has been published. ErrDisabled — a reload +// switched the store off after the settings snapshot was read — publishes the +// window un-deduped, as records under the other setting would have been. +func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, table string, recs []pendingRecord, keyed []int) *requestAbort { lease := h.DedupeLease if lease <= 0 { lease = dedupe.DefaultLease } - claims, err := dd.Reserve(ctx, []dedupe.Key{key}, lease) + keys := make([]dedupe.Key, len(keyed)) + for j, i := range keyed { + keys[j] = *recs[i].key + } + claims, err := dd.Reserve(ctx, keys, lease) switch { case errors.Is(err, dedupe.ErrDisabled): // The counter carries the signal (a burst is a reload; a steady rate // is the store and settings out of step), so the line is Debug rather // than a WARN per record. - dedupeDisabledCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", key.Table))) - slog.DebugContext(ctx, "dedupe switched off mid-reload; publishing without idempotency", "event_id", key.ID, "table", key.Table) - return nil, false, nil + dedupeDisabledCounter.Add(ctx, int64(len(keys)), metric.WithAttributes(attribute.String("table", table))) + slog.DebugContext(ctx, "dedupe switched off mid-reload; publishing without idempotency", "records", len(keys), "table", table) + return nil + case errors.Is(err, dedupe.ErrUnavailable): + slog.WarnContext(ctx, "dedupe store unavailable", "error", err, "table", table) + return &requestAbort{Status: http.StatusServiceUnavailable, Message: "dedupe store unavailable", RetryAfter: "5"} case err != nil: - slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "event_id", key.ID, "table", key.Table) - return nil, false, &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} - } - switch claims[0].Status { - case dedupe.Duplicate: - slog.InfoContext(ctx, "duplicate event skipped", "event_id", key.ID, "table", key.Table) - return nil, true, nil - case dedupe.InFlight: - slog.InfoContext(ctx, "event id in flight in another request", "event_id", key.ID, "table", key.Table) - return nil, false, &requestAbort{ + slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "table", table) + return &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} + } + var held *dedupe.Key + for j, i := range keyed { + recs[i].claim = claims[j] + switch claims[j].Status { + case dedupe.Duplicate: + recs[i].duplicate = true + slog.InfoContext(ctx, "duplicate event skipped", "event_id", keys[j].ID, "table", table) + case dedupe.InFlight: + if held == nil { + held = &keys[j] + } + case dedupe.Claimed: + } + } + if held != nil { + releaseClaims(ctx, dd, claimedIn(recs)) + slog.InfoContext(ctx, "event id in flight in another request", "event_id", held.ID, "table", table) + return &requestAbort{ Status: http.StatusServiceUnavailable, Message: "a request with the same dedupe id is in flight", RetryAfter: strconv.Itoa(int(math.Ceil(lease.Seconds()))), } - case dedupe.Claimed: } - return claims, false, nil + return nil +} + +// publishFailed settles a window whose publish failed at recs[k] and returns +// the abort. The records before k are queued, so their ids are committed. A +// definite failure — ErrQueueFull, the broker refused the event — releases k's +// id and the rest, so the client's retry publishes them (#384). Any other +// failure may have stored the event before failing, so k's claim is left to +// lapse with its lease instead: a retry before then answers in-flight, and one +// after republishes under the same idempotency key, which the queue drops if +// the first copy landed. The records after k were never sent and are released. +func (h *IngestHandler) publishFailed(ctx context.Context, dd dedupe.Deduplicator, topic mq.Topic, recs []pendingRecord, k int, err error) *requestAbort { + definite := errors.Is(err, mq.ErrQueueFull) + commitClaims(ctx, dd, claimedIn(recs[:k]), topic.Table) + after := k + 1 + if definite { + after = k + } + releaseClaims(ctx, dd, claimedIn(recs[after:])) + if definite { + slog.WarnContext(ctx, "ingest queue is full", "tenant", topic.Tenant, "error", err, "table", topic.Table, "scope", topic.Scope) + return &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} + } + slog.ErrorContext(ctx, "failed to publish to the ingest queue", "tenant", topic.Tenant, "error", err, "table", topic.Table, "scope", topic.Scope) + return &requestAbort{Status: http.StatusInternalServerError, Message: "publish failed"} +} + +// claimedIn is the Claimed claims among recs. +func claimedIn(recs []pendingRecord) []dedupe.Claim { + var out []dedupe.Claim + for i := range recs { + if recs[i].claim.Status == dedupe.Claimed { + out = append(out, recs[i].claim) + } + } + return out } -// commitClaims makes a published record's id a duplicate. A failure does not -// fail the record — it is in the queue — so it is logged and counted, and -// the claim lapses after its lease. +// commitClaims makes published records' ids duplicates. A failure does not +// fail the records — they are in the queue — so it is logged and counted, and +// the claims lapse after their lease. func commitClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim, table string) { if len(claims) == 0 { return } - // The record is queued whatever the request's context does next. + // The records are queued whatever the request's context does next. err := dd.Commit(context.WithoutCancel(ctx), claims, 0) switch { case err == nil, errors.Is(err, dedupe.ErrDisabled): default: - dedupeCommitFailedCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) - slog.ErrorContext(ctx, "dedupe commit failed after publish; the id lapses with its lease", "error", err, "table", table) + dedupeCommitFailedCounter.Add(ctx, int64(len(claims)), metric.WithAttributes(attribute.String("table", table))) + slog.ErrorContext(ctx, "dedupe commit failed after publish; the ids lapse with their lease", "error", err, "table", table, "records", len(claims)) } } -// releaseClaims gives claims back after a failed publish. A failure is only -// logged: the claim lapses with its lease either way. +// releaseClaims gives back claims whose records were not published. A failure +// is only logged: the claims lapse with their lease either way. func releaseClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim) { if len(claims) == 0 { return } if err := dd.Release(context.WithoutCancel(ctx), claims); err != nil && !errors.Is(err, dedupe.ErrDisabled) { - slog.WarnContext(ctx, "dedupe release failed; the id lapses with its lease", "error", err) + slog.WarnContext(ctx, "dedupe release failed; the ids lapse with their lease", "error", err) } } diff --git a/internal/api/ingest_seams.go b/internal/api/ingest_seams.go index d95fdb58f..210f2ba59 100644 --- a/internal/api/ingest_seams.go +++ b/internal/api/ingest_seams.go @@ -20,7 +20,7 @@ import ( // return would invite a caller to change that. // // The two are one interface because they are one contract — "what this schema -// says about this record" — evaluated at two points in processRecord that must +// says about this record" — evaluated at two points in prepareRecord that must // stay apart: the insert-check block sits between them deliberately, so checks // keep pre-#372 semantics. type RecordValidator interface { @@ -55,7 +55,7 @@ func (h *IngestHandler) validator() RecordValidator { // InsertChecker decides whether a record's value satisfies a policy check // clause. Matches answers the scalar `_eq` form (the required value), InSet the // `_in` form (set membership). It never sees a record as a whole: the -// auto-injection of a missing check value stays in processRecord, where the +// auto-injection of a missing check value stays in prepareRecord, where the // ordering against validation and canonicalization is load-bearing. type InsertChecker interface { Matches(actual, required any) bool diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index f08d2eb02..579933339 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -660,7 +660,7 @@ func TestIngest_Policy_CheckIn_Absent_FailsClosed(t *testing.T) { // TestIngest_Policy_CheckIn_AbsentClaim_FailsClosed locks the typed-nil []any // path behind an _in check: when the claim itself is absent, resolveInValues -// returns a typed-nil []any, which must still assert as []any in processRecord +// returns a typed-nil []any, which must still assert as []any in prepareRecord // (entering the membership branch) so the column is rejected — never treated as a // scalar _eq value and auto-injected. The sibling _Absent test omits the column // with the claim present; this one drops the claim too. Guards #224 fail-closed. @@ -1825,8 +1825,8 @@ func TestIngest_JSONArray_SyntaxError_Fatal(t *testing.T) { h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) // A structural syntax error desyncs the decoder — the whole request fails - // (400), unlike a per-element type error. The leading good element may have - // already published (at-least-once on retry). + // (400), unlike a per-element type error. The leading good element is still + // in the open window, which is dropped unpublished. req := rawIngestRequest(t, "clicks", "application/json", `[{"page":"/a"}, {bad]`) w := httptest.NewRecorder() h.Handle(w, withTenant(req)) @@ -1834,7 +1834,7 @@ func TestIngest_JSONArray_SyntaxError_Fatal(t *testing.T) { assert.Equal(t, http.StatusBadRequest, w.Code) assert.Contains(t, w.Body.String(), "invalid json") testutil.AssertJSONErrorResponse(t, w) - assert.Len(t, pub.Messages, 1) // the leading record published before the abort + assert.Empty(t, pub.Messages) } func TestIngest_JSONArray_Truncated_Fatal(t *testing.T) { @@ -2347,7 +2347,7 @@ func TestIngest_Dedup_DisabledMidReload(t *testing.T) { // discovery.Validate accepts `{}` here because every column is nullable or // defaulted. I previously asserted this path was unreachable, having tested only // against a schema with a required column; it is not. -func TestProcessRecord_UnresolvedInsertSideAborts(t *testing.T) { +func TestPrepareRecord_UnresolvedInsertSideAborts(t *testing.T) { t.Parallel() schema := &discovery.TableSchema{ Name: "loose", @@ -2371,11 +2371,10 @@ func TestProcessRecord_UnresolvedInsertSideAborts(t *testing.T) { require.NoError(t, discovery.Validate(schema, map[string]any{}), "all-nullable/defaulted columns accept an empty record — this is what makes the read reachable") - dup, reject, abort := h.processRecord( + rec, abort := h.prepareRecord( context.Background(), testStore, "loose", "", schema, selectResolved, "viewer", map[string]any{}, time.Now(), nil) - assert.False(t, dup) - assert.Nil(t, reject, "a request-scoped condition must not be reported per record") + assert.Nil(t, rec.reject, "a request-scoped condition must not be reported per record") require.NotNil(t, abort, "an unresolved insert side must abort the request") assert.Equal(t, http.StatusForbidden, abort.Status) assert.Empty(t, abort.RetryAfter, "not a transient condition — retrying cannot help") @@ -2749,43 +2748,79 @@ func dedupHandler(t *testing.T, pub *testutil.MockPublisher, dedup dedupe.Dedupl return h } -// #384: a publish that fails gives the id back, so the retry the 503 asks for -// is published rather than skipped as a duplicate of a record that never -// reached the queue. +// #384: a publish the queue refused gives the id back, so the retry the 503 +// asks for is published rather than skipped as a duplicate of a record that +// never reached the queue. func TestIngest_Dedup_FailedPublishReleasesTheID(t *testing.T) { t.Parallel() - tests := []struct { - name string - err error - status int - }{ - {"backpressure", fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull), http.StatusServiceUnavailable}, - {"other failure", errors.New("connection reset"), http.StatusInternalServerError}, - } - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - pub := &testutil.MockPublisher{Err: tt.err} - dedup := testutil.NewMockDeduplicator() - h := dedupHandler(t, pub, dedup, false) - body := map[string]any{"page": "/home", "event_id": "e1"} + pub := &testutil.MockPublisher{Err: fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull)} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + body := map[string]any{"page": "/home", "event_id": "e1"} - w := httptest.NewRecorder() - h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) - require.Equal(t, tt.status, w.Code) - assert.False(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "released, not left to lapse") - - pub.Err = nil - w = httptest.NewRecorder() - h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) - require.Equal(t, http.StatusOK, w.Code) - assert.Contains(t, w.Body.String(), `"ok":true`, "the retry is published, not a duplicate") - assert.Len(t, pub.Published(), 1) - - w = httptest.NewRecorder() - h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) - assert.Contains(t, w.Body.String(), `"duplicate":true`, "and committed once published") - }) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "30", w.Header().Get("Retry-After")) + assert.False(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "released, not left to lapse") + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusOK, w.Code) + assert.Contains(t, w.Body.String(), `"ok":true`, "the retry is published, not a duplicate") + assert.Len(t, pub.Published(), 1) + + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + assert.Contains(t, w.Body.String(), `"duplicate":true`, "and committed once published") +} + +// A publish whose outcome is unknown may have stored the event, so its claim +// is neither released nor committed: it lapses with the lease, a retry before +// then answers in-flight, and the idempotency key covers one after. +func TestIngest_Dedup_UncertainPublishLeavesTheClaim(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: context.DeadlineExceeded} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + body := map[string]any{"page": "/home", "event_id": "e1"} + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + require.Equal(t, http.StatusInternalServerError, w.Code) + assert.Contains(t, w.Body.String(), "publish failed") + assert.True(t, dedup.Pending(dedupe.Key{Table: "clicks", ID: "e1"}), "left to lapse") + assert.Empty(t, dedup.Released) + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "30", w.Header().Get("Retry-After")) + assert.Empty(t, pub.Published()) +} + +// A claimed record is published under its idempotency key; an un-deduped one +// carries none, so a producer's repeated ids are not dropped by the queue. +func TestIngest_Dedup_PublishCarriesTheIdempotencyKey(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, testutil.NewMockDeduplicator(), false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", + jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), + jsonLine(t, map[string]any{"page": "/b"}), + ))) + require.Equal(t, http.StatusOK, w.Code) + msgs := pub.Published() + require.Len(t, msgs, 2) + + want := mq.Headers{} + mq.WithIdempotencyKey(dedupe.IdempotencyKey(testStore.Tenant(), dedupe.Key{Table: "clicks", ID: "e1"}))(want) + for k, v := range want { + assert.Equal(t, v, msgs[0].Headers[k]) + assert.NotContains(t, msgs[1].Headers, k) } } @@ -2808,8 +2843,9 @@ func TestIngest_NDJSON_Dedup_PublishFailureMidBatch(t *testing.T) { w := httptest.NewRecorder() h.Handle(w, withTenant(batch())) require.Equal(t, http.StatusServiceUnavailable, w.Code) - require.Len(t, dedup.Released, 1) + require.Len(t, dedup.Released, 2, "the failing record and the rest of its window") assert.Equal(t, dedupe.Key{Table: "clicks", ID: "e2"}, dedup.Released[0].Key) + assert.Equal(t, dedupe.Key{Table: "clicks", ID: "e3"}, dedup.Released[1].Key) pub.Err = nil w = httptest.NewRecorder() diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go new file mode 100644 index 000000000..21f2e0752 --- /dev/null +++ b/internal/api/ingest_window_test.go @@ -0,0 +1,435 @@ +package api + +import ( + "context" + "errors" + "fmt" + "net/http" + "net/http/httptest" + "strings" + "sync" + "testing" + "time" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/testutil" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// eventLines is n NDJSON clicks records with ids e1..en. +func eventLines(t *testing.T, n int) []string { + t.Helper() + lines := make([]string, n) + for i := range n { + lines[i] = jsonLine(t, map[string]any{"page": "/p", "event_id": fmt.Sprintf("e%d", i+1)}) + } + return lines +} + +func clickKey(i int) dedupe.Key { return dedupe.Key{Table: "clicks", ID: fmt.Sprintf("e%d", i)} } + +// A batch is reserved, published and committed a window at a time: one +// Reserve and one Commit per window, whatever the batch size. +func TestIngest_Windows_OneReserveAndCommitPerWindow(t *testing.T) { + t.Parallel() + for _, n := range []int{1, 255, 256, 257, 600} { + t.Run(fmt.Sprint(n), func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, n)...))) + require.Equal(t, http.StatusOK, w.Code) + assert.Equal(t, n, decodeBatchResult(t, w).Succeeded) + windows := (n + ingestWindow - 1) / ingestWindow + assert.Equal(t, windows, dedup.Reserves) + assert.Equal(t, windows, dedup.Commits) + assert.Len(t, pub.Published(), n) + assert.True(t, dedup.Committed(clickKey(n))) + }) + } +} + +// A publish failing at record k settles its window: the records before k are +// committed, k is released when the queue refused it and left to lapse when +// the outcome is unknown, the rest of the window is released, and later +// windows are never reserved. A whole-batch retry after a refusal publishes +// every record exactly once. +func TestIngest_Windows_PublishFailureAtK(t *testing.T) { + t.Parallel() + const n = 600 + refused := fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull) + tests := []struct { + name string + k int + err error + status int + }{ + {"refused first record", 1, refused, http.StatusServiceUnavailable}, + {"refused mid first window", 100, refused, http.StatusServiceUnavailable}, + {"refused last of first window", 256, refused, http.StatusServiceUnavailable}, + {"refused first of second window", 257, refused, http.StatusServiceUnavailable}, + {"refused mid last window", 590, refused, http.StatusServiceUnavailable}, + {"uncertain mid first window", 100, context.DeadlineExceeded, http.StatusInternalServerError}, + {"uncertain mid second window", 400, context.DeadlineExceeded, http.StatusInternalServerError}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{Err: tt.err, ErrAfter: tt.k - 1} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + lines := eventLines(t, n) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", lines...))) + require.Equal(t, tt.status, w.Code) + assert.Len(t, pub.Published(), tt.k-1) + + windowEnd := min((tt.k-1)/ingestWindow*ingestWindow+ingestWindow, n) + definite := errors.Is(tt.err, mq.ErrQueueFull) + var released []dedupe.Key + for _, c := range dedup.Released { + released = append(released, c.Key) + } + var wantReleased []dedupe.Key + for i := tt.k; i <= windowEnd; i++ { + if i > tt.k || definite { + wantReleased = append(wantReleased, clickKey(i)) + } + } + assert.Equal(t, wantReleased, released) + if tt.k > 1 { + assert.True(t, dedup.Committed(clickKey(1))) + assert.True(t, dedup.Committed(clickKey(tt.k-1)), "published before the failure") + } + assert.False(t, dedup.Committed(clickKey(tt.k))) + assert.Equal(t, !definite, dedup.Pending(clickKey(tt.k)), "an uncertain publish leaves its claim to lapse") + if windowEnd < n { + assert.False(t, dedup.Pending(clickKey(windowEnd+1)), "a later window is never reserved") + } + if !definite { + return + } + + pub.Err = nil + w = httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", lines...))) + require.Equal(t, http.StatusOK, w.Code) + resp := decodeBatchResult(t, w) + assert.Equal(t, tt.k-1, resp.Duplicates) + assert.Equal(t, n-(tt.k-1), resp.Succeeded) + assert.Len(t, pub.Published(), n, "every record exactly once") + }) + } +} + +// A dedupe store that cannot answer now fails the request with 503 and a +// short Retry-After, which the SDK retries — not the 500 of a broken store. +// Earlier windows stay published and committed. +func TestIngest_Dedup_UnavailableIs503(t *testing.T) { + t.Parallel() + notOpen := dedupe.NewManaged(func() (dedupe.Deduplicator, error) { return nil, errors.New("disk gone") }) + require.Error(t, notOpen.Apply(true)) + throttled := testutil.NewMockDeduplicator() + throttled.Err = fmt.Errorf("%w: throttled", dedupe.ErrUnavailable) + secondWindow := testutil.NewMockDeduplicator() + secondWindow.Err, secondWindow.ErrAfter = throttled.Err, 1 + + tests := []struct { + name string + dedup dedupe.Deduplicator + n int + published int + }{ + {"store not open", notOpen, 1, 0}, + {"backend throttled", throttled, 3, 0}, + {"second window throttled", secondWindow, ingestWindow + 1, ingestWindow}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, tt.dedup, false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, tt.n)...))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "5", w.Header().Get("Retry-After")) + assert.Contains(t, w.Body.String(), "dedupe store unavailable") + assert.Len(t, pub.Published(), tt.published) + }) + } + t.Run("single object", func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + h := dedupHandler(t, pub, throttled, false) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/", "event_id": "e1"}))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "5", w.Header().Get("Retry-After")) + assert.Empty(t, pub.Published()) + }) +} + +// One id held by another request stops its window before anything in it is +// published and gives back the window's other claims; windows before it stay +// committed. +func TestIngest_Windows_InFlightReleasesTheWindow(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + held := ingestWindow + 2 + dedup.Hold(clickKey(held)) + h := dedupHandler(t, pub, dedup, false) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, ingestWindow+3)...))) + assert.Equal(t, http.StatusServiceUnavailable, w.Code) + assert.Equal(t, "30", w.Header().Get("Retry-After")) + assert.Len(t, pub.Published(), ingestWindow) + assert.True(t, dedup.Committed(clickKey(ingestWindow))) + for _, i := range []int{ingestWindow + 1, ingestWindow + 3} { + assert.False(t, dedup.Pending(clickKey(i)), "e%d released", i) + } + assert.True(t, dedup.Pending(clickKey(held)), "the other request's claim is untouched") +} + +// Rejects, duplicates and repeats keep their places in the results across +// windows, over both batch formats. +func TestIngest_Windows_OutcomesStayInOrder(t *testing.T) { + t.Parallel() + records := []map[string]any{ + {"page": "/a", "event_id": "e1"}, + {"page": "/b", "event_id": "e1"}, // repeat inside one window + {"page": "/c", "nope": 1}, // reject + {"page": "/d", "event_id": "e2"}, + {"page": "/e", "event_id": "e1"}, // repeat across windows + {"page": "/f"}, // no id: published un-deduped + } + requests := map[string]func() *http.Request{ + "ndjson": func() *http.Request { + lines := make([]string, len(records)) + for i, r := range records { + lines[i] = jsonLine(t, r) + } + return ndjsonRequest(t, "clicks", lines...) + }, + "json array": func() *http.Request { return ingestRequest(t, "clicks", records) }, + } + for name, req := range requests { + t.Run(name, func(t *testing.T) { + t.Parallel() + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + h := dedupHandler(t, pub, dedup, false) + h.window = 3 + + w := httptest.NewRecorder() + h.Handle(w, withTenant(req())) + require.Equal(t, http.StatusOK, w.Code) + resp := decodeBatchResult(t, w) + assert.Equal(t, []recordResult{ + {Index: 1, Ok: true}, + {Index: 2, Duplicate: true}, + {Index: 3, Error: resp.Results[2].Error}, + {Index: 4, Ok: true}, + {Index: 5, Duplicate: true}, + {Index: 6, Ok: true}, + }, resp.Results) + assert.NotEmpty(t, resp.Results[2].Error) + assert.Equal(t, 6, resp.Total) + assert.Len(t, pub.Published(), 3) + assert.Equal(t, 2, dedup.Reserves) + }) + } +} + +// The embedded queue must remember an idempotency key for at least a lease: +// the retry of an uncertain publish lands after the lease, and only the queue's +// duplicate window drops its second copy. +func TestIngest_DedupeLeaseFitsTheDuplicateWindow(t *testing.T) { + t.Parallel() + assert.LessOrEqual(t, dedupe.DefaultLease, mq.EmbeddedDuplicateWindow) +} + +// faultyPublisher publishes through a real broker and fails the calls fail +// picks: before sending (the queue refused it) or after (the outcome unknown +// to the caller, though the event is stored). +type faultyPublisher struct { + mq.Publisher + mu sync.Mutex + calls int + fail func(call int) (sendFirst bool, err error) +} + +func (p *faultyPublisher) Publish(ctx context.Context, topic mq.Topic, data []byte, opts ...mq.PublishOpt) error { + p.mu.Lock() + p.calls++ + sendFirst, err := p.fail(p.calls) + p.mu.Unlock() + if err == nil || sendFirst { + if pubErr := p.Publisher.Publish(ctx, topic, data, opts...); pubErr != nil { + return pubErr + } + } + return err +} + +// realPipeline is an ingest handler over the embedded broker and Pebble +// store, with pub's faults in front of the broker, and a count of the events +// in the tenant's queue. +func realPipeline(t *testing.T, fail func(call int) (bool, error)) (*IngestHandler, func() int) { + t.Helper() + broker, err := mq.NewEmbedded(t.TempDir()) + require.NoError(t, err) + t.Cleanup(func() { _ = broker.Close() }) + require.NoError(t, broker.SetMaxBytes(t.Context(), testStore.Tenant(), 64<<20)) + store := dedupe.NewEmbedded(t.TempDir()).Tenant(testStore.Tenant()) + require.NoError(t, store.Apply(true)) + t.Cleanup(func() { _ = store.Close() }) + + h := dedupHandler(t, nil, store, false) + h.Publisher = &faultyPublisher{Publisher: broker, fail: fail} + count := func() int { + n := 0 + require.NoError(t, broker.ReplaySince(t.Context(), mq.Topic{Tenant: testStore.Tenant(), Table: "clicks"}, time.Time{}, + func([]byte) bool { n++; return true })) + return n + } + return h, count +} + +// #384 end to end: a publish the queue refused, then the client's retry, ends +// in exactly one event in the queue — and a later retry is a duplicate. +func TestIngest_Dedup_FailedPublishThenRetryIsOneEvent(t *testing.T) { + t.Parallel() + h, count := realPipeline(t, func(call int) (bool, error) { + if call == 1 { + return false, fmt.Errorf("%w: maximum bytes exceeded", mq.ErrQueueFull) + } + return false, nil + }) + body := map[string]any{"page": "/home", "event_id": "e1"} + codes := make([]int, 3) + for i := range codes { + w := httptest.NewRecorder() + h.Handle(w, withTenant(ingestRequest(t, "clicks", body))) + codes[i] = w.Code + if i == 2 { + assert.Contains(t, w.Body.String(), `"duplicate":true`) + } + } + assert.Equal(t, []int{http.StatusServiceUnavailable, http.StatusOK, http.StatusOK}, codes) + assert.Equal(t, 1, count()) +} + +// A publish that stored the event but reported a failure, then the client's +// retry: in-flight until the lease lapses, then republished under the same +// idempotency key, which the queue drops — one event, and the id committed. +func TestIngest_Dedup_UncertainPublishThenRetryIsOneEvent(t *testing.T) { + t.Parallel() + h, count := realPipeline(t, func(call int) (bool, error) { + if call == 1 { + return true, context.DeadlineExceeded + } + return false, nil + }) + h.DedupeLease = 300 * time.Millisecond + lines := eventLines(t, 3) + send := func() *httptest.ResponseRecorder { + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", lines...))) + return w + } + + w := send() + require.Equal(t, http.StatusInternalServerError, w.Code) + w = send() + require.Equal(t, http.StatusServiceUnavailable, w.Code, "the uncertain claim is still held") + assert.Equal(t, "1", w.Header().Get("Retry-After")) + + var last *httptest.ResponseRecorder + require.Eventually(t, func() bool { + last = send() + return last.Code == http.StatusOK + }, 5*time.Second, 50*time.Millisecond) + assert.Equal(t, 3, decodeBatchResult(t, last).Succeeded, "the lapsed claim is claimed again and republished") + assert.Equal(t, 3, count(), "the republished e1 was dropped by the queue") + + w = send() + require.Equal(t, http.StatusOK, w.Code) + assert.Equal(t, 3, decodeBatchResult(t, w).Duplicates) +} + +// countingDedup counts the Commits that reach a store: on Pebble each is one +// fsync. +type countingDedup struct { + dedupe.Deduplicator + mu sync.Mutex + commits int +} + +func (c *countingDedup) Commit(ctx context.Context, claims []dedupe.Claim, retention time.Duration) error { + c.mu.Lock() + c.commits++ + c.mu.Unlock() + return c.Deduplicator.Commit(ctx, claims, retention) +} + +// pebbleBatchHandler is a handler over a real Pebble store behind a Commit +// counter, publishing to a mock queue. +func pebbleBatchHandler(tb testing.TB, window int) (*IngestHandler, *countingDedup) { + tb.Helper() + store := dedupe.NewEmbedded(tb.TempDir()).Tenant(testStore.Tenant()) + require.NoError(tb, store.Apply(true)) + tb.Cleanup(func() { _ = store.Close() }) + counted := &countingDedup{Deduplicator: store} + h := NewIngestHandler(fixedRegistry(testRegistry(tb)), &testutil.MockPublisher{}) + h.Dedup = staticDedup(counted) + h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.window = window + return h, counted +} + +// Windows cut the per-record fsyncs on Pebble: a 1,000-record batch commits in +// four syncs rather than a thousand. +func TestIngest_Windows_OneSyncPerWindowOnPebble(t *testing.T) { + t.Parallel() + h, counted := pebbleBatchHandler(t, 0) + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", eventLines(t, 1000)...))) + require.Equal(t, http.StatusOK, w.Code) + assert.Equal(t, 4, counted.commits) +} + +// BenchmarkIngest_DedupBatchOnPebble compares a 1,000-record batch committed +// per record (window 1, the pre-window behavior) with the default window. +func BenchmarkIngest_DedupBatchOnPebble(b *testing.B) { + for _, window := range []int{1, ingestWindow} { + b.Run(fmt.Sprintf("window=%d", window), func(b *testing.B) { + h, counted := pebbleBatchHandler(b, window) + var body strings.Builder + iter := 0 + for b.Loop() { + iter++ + body.Reset() + for i := range 1000 { + fmt.Fprintf(&body, `{"page":"/p","event_id":"%d-%d"}`+"\n", iter, i) + } + req := httptest.NewRequestWithContext(context.Background(), http.MethodPost, "/v1/ingest?table=clicks", strings.NewReader(body.String())) + req.Header.Set("Content-Type", "application/x-ndjson") + w := httptest.NewRecorder() + h.Handle(w, withTenant(req)) + if w.Code != http.StatusOK { + b.Fatalf("status %d", w.Code) + } + } + b.ReportMetric(float64(counted.commits)/float64(iter), "syncs/op") + }) + } +} diff --git a/internal/dedupe/key.go b/internal/dedupe/key.go index 25ead7c26..892e7c3f7 100644 --- a/internal/dedupe/key.go +++ b/internal/dedupe/key.go @@ -2,6 +2,7 @@ package dedupe import ( "crypto/sha256" + "encoding/hex" "errors" "fmt" "strings" @@ -69,3 +70,11 @@ func AppendKey(dst, prefix []byte, k Key) []byte { } return append(dst, k.ID...) } + +// IdempotencyKey is k's message id for the queue under tenant id: the first +// 128 bits of the stored key's SHA-256, in hex, so a republished record is +// recognised without its id riding in a header verbatim. +func IdempotencyKey(id tenant.ID, k Key) string { + sum := sha256.Sum256(AppendKey(nil, KeyPrefix(id), k)) + return hex.EncodeToString(sum[:16]) +} diff --git a/internal/dedupe/key_test.go b/internal/dedupe/key_test.go new file mode 100644 index 000000000..2f2981e22 --- /dev/null +++ b/internal/dedupe/key_test.go @@ -0,0 +1,33 @@ +package dedupe_test + +import ( + "strings" + "testing" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/stretchr/testify/assert" +) + +// The idempotency key is 32 hex characters, stable for one tenant, table and +// id, and different when any of the three differs — ids too long to store +// verbatim included. +func TestIdempotencyKey(t *testing.T) { + t.Parallel() + long := strings.Repeat("x", dedupe.MaxIDBytes+1) + base := dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: "e1"}) + assert.Regexp(t, `^[0-9a-f]{32}$`, base) + assert.Equal(t, base, dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: "e1"})) + + others := []string{ + dedupe.IdempotencyKey("globex", dedupe.Key{Table: "clicks", ID: "e1"}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "views", ID: "e1"}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: "e2"}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: long}), + dedupe.IdempotencyKey("acme", dedupe.Key{Table: "clicks", ID: long + "y"}), + } + seen := map[string]bool{base: true} + for _, k := range others { + assert.False(t, seen[k], "collision: %s", k) + seen[k] = true + } +} diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 830be9003..7ccc6e3e9 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -306,17 +306,24 @@ func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { } } +// EmbeddedDuplicateWindow is how long an ingest queue remembers a +// WithIdempotencyKey key. A dedupe lease must not exceed it: a claim left to +// lapse after an uncertain publish is republished once the lease ends, and +// only this window drops that second copy. +const EmbeddedDuplicateWindow = 2 * time.Minute + // ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard // append-only log; the Active Sweeper handles message purging. MaxBytes caps // the tenant's share of the disk. DiscardNew rejects new messages when full, // propagating backpressure to the upstream API — for this tenant alone. func ingestStreamConfig(id tenant.ID, maxBytes int64) jetstream.StreamConfig { return jetstream.StreamConfig{ - Name: ingestStreamName(id), - Subjects: []string{tenantSubjects(ingestPrefix, id)}, - Retention: jetstream.LimitsPolicy, - MaxBytes: maxBytes, - Discard: jetstream.DiscardNew, + Name: ingestStreamName(id), + Subjects: []string{tenantSubjects(ingestPrefix, id)}, + Retention: jetstream.LimitsPolicy, + MaxBytes: maxBytes, + Discard: jetstream.DiscardNew, + Duplicates: EmbeddedDuplicateWindow, } } diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 6e87ff7b7..004938d02 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -141,6 +141,34 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { assert.Equal(t, []byte("x"), raw.Data) } +// A repeated idempotency key inside the duplicate window is dropped as a +// success, so an uncertain publish can be republished safely; a queue made +// before the window was set gets it on its next budget apply. +func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + old := ingestStreamConfig(tenant.Default, testBudget) + old.Duplicates = 0 + _, err := e.js.CreateStream(ctx, old) + require.NoError(t, err) + require.NoError(t, e.SetMaxBytes(ctx, tenant.Default, testBudget)) + require.Equal(t, EmbeddedDuplicateWindow, streamConfig(t, e, "INGEST_0").Duplicates) + + topic := Topic{Tenant: tenant.Default, Table: "t"} + require.NoError(t, e.Publish(ctx, topic, []byte("a"), WithIdempotencyKey("k1"))) + require.NoError(t, e.Publish(ctx, topic, []byte("a again"), WithIdempotencyKey("k1")), "a repeat is a success") + require.NoError(t, e.Publish(ctx, topic, []byte("b"), WithIdempotencyKey("k2"))) + require.NoError(t, e.Publish(ctx, topic, []byte("c"))) + + var got []string + require.NoError(t, e.ReplaySince(ctx, topic, time.Time{}, func(data []byte) bool { + got = append(got, string(data)) + return true + })) + assert.Equal(t, []string{"a", "b", "c"}, got) +} + // A tenant's first budget opens its queue: an ingest stream holding its // subjects alone at the budget, refusing when full, and a dead-letter stream // at a tenth of it, dropping its oldest when full. No other tenant gets one. @@ -155,6 +183,7 @@ func TestEmbeddedNATS_SetMaxBytes_OpensTheTenantsQueue(t *testing.T) { assert.Equal(t, []string{"ingest.acme.>"}, ingest.Subjects) assert.Equal(t, int64(testBudget), ingest.MaxBytes) assert.Equal(t, jetstream.DiscardNew, ingest.Discard) + assert.Equal(t, EmbeddedDuplicateWindow, ingest.Duplicates) dlq := streamConfig(t, e, "DLQ_acme") assert.Equal(t, []string{"dlq.acme.>"}, dlq.Subjects) assert.Equal(t, int64(testBudget)/10, dlq.MaxBytes) diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 3f1c45c1f..97eac3654 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -141,6 +141,20 @@ func WithHeader(key, value string) PublishOpt { } } +// idempotencyHeader carries WithIdempotencyKey's key: JetStream's own +// message-id header, which the stream deduplicates on. +const idempotencyHeader = "Nats-Msg-Id" + +// WithIdempotencyKey marks a publish with key: a second publish carrying the +// same key within the queue's duplicate window is dropped by the broker and +// reported as success, so republishing an event whose first publish had an +// unknown outcome stores it once. +func WithIdempotencyKey(key string) PublishOpt { + return func(h Headers) { + h.Set(idempotencyHeader, key) + } +} + // ErrQueueFull is returned by Publisher.Publish when the topic's tenant's // ingest queue refuses new events — it is at its byte budget, or the tenant // has no queue open yet — the backpressure signal the API turns into a 503 diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index cdb1e33e7..624fe0efd 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -118,11 +118,15 @@ type MockDeduplicator struct { committed map[dedupe.Key]bool pending map[dedupe.Key]string tokens int - // Err, if set, fails Reserve; CommitErr and ReleaseErr fail their phase. + // Err, if set, fails Reserve — after ErrAfter calls have succeeded; + // CommitErr and ReleaseErr fail their phase. Err error + ErrAfter int CommitErr error ReleaseErr error Released []dedupe.Claim // every claim Release was given + // Calls to each phase, for tests that count round trips. + Reserves, Commits int } var _ dedupe.Deduplicator = (*MockDeduplicator)(nil) @@ -131,16 +135,21 @@ func NewMockDeduplicator() *MockDeduplicator { return &MockDeduplicator{committed: map[dedupe.Key]bool{}, pending: map[dedupe.Key]string{}} } +// Reserve answers Duplicate for a key repeated in one call, as Managed does. func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time.Duration) ([]dedupe.Claim, error) { - if m.Err != nil { - return nil, m.Err - } m.mu.Lock() defer m.mu.Unlock() + m.Reserves++ + if m.Err != nil && m.Reserves > m.ErrAfter { + return nil, m.Err + } claims := make([]dedupe.Claim, 0, len(keys)) + seen := make(map[dedupe.Key]bool, len(keys)) for _, k := range keys { + repeat := seen[k] + seen[k] = true switch { - case m.committed[k]: + case repeat, m.committed[k]: claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.Duplicate}) case m.pending[k] != "": claims = append(claims, dedupe.Claim{Key: k, Status: dedupe.InFlight}) @@ -155,11 +164,12 @@ func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time. } func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ time.Duration) error { + m.mu.Lock() + defer m.mu.Unlock() + m.Commits++ if m.CommitErr != nil { return m.CommitErr } - m.mu.Lock() - defer m.mu.Unlock() for _, c := range claims { if c.Status == dedupe.Claimed { m.committed[c.Key] = true @@ -192,6 +202,13 @@ func (m *MockDeduplicator) Hold(k dedupe.Key) { m.pending[k] = "held" } +// Committed reports whether k was committed. +func (m *MockDeduplicator) Committed(k dedupe.Key) bool { + m.mu.Lock() + defer m.mu.Unlock() + return m.committed[k] +} + // Pending reports whether k is claimed and neither committed nor released. func (m *MockDeduplicator) Pending(k dedupe.Key) bool { m.mu.Lock() From 71ab82482ca057dac5a47b0ae2fd6efbd3927f39 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:48:41 -0400 Subject: [PATCH 02/22] fix(ingest): qualify the duplicate-window claims; steadier tests The idempotency key drops a retry only within two minutes of the first publish, and a 200 with a failed commit is counted, not committed: the docs now say so. The uncertain-publish test uses a 2s lease so a stall cannot lapse the claim early, and the mq test pins the explicit duplicate window rather than the server's matching default. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 4 ++-- docs/src/content/docs/settings-directory.mdx | 2 +- internal/api/ingest_window_test.go | 6 +++--- internal/mq/embedded_test.go | 6 ++++-- 7 files changed, 13 insertions(+), 11 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 041ef99e4..f0dab22eb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and the retry after the lease is dropped by the queue if the first copy was stored. A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 11f9f1f07..dd84c32e1 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -274,7 +274,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now (not open, or a remote backend throttled or unreachable); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored (it remembers the key for two minutes), so the retry never stores a second copy. | +| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue remembers the key for two minutes after the first publish, so a retry inside that window stores no second copy (the SDK's, after the 30-second `Retry-After`, lands inside it); a later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index eb660c17d..e50b269fd 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -80,7 +80,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. -- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates and encodes each record, and runs the records in windows of up to 256 (`ingestWindow`) through three phases: one dedupe `Reserve` for the window's ids, the publishes in record order (a deduped record under `mq.WithIdempotencyKey`, keyed by `dedupe.IdempotencyKey`), and one `Commit` of the published ids — a window is the unit of a dedupe round trip and of Pebble's commit `fsync`. An id another request holds answers `503` with the lease as `Retry-After`, a store that cannot answer (`dedupe.ErrUnavailable`) `503` with `Retry-After: 5`; a publish that fails at a record commits the ones before it and releases the rest, except that a failure other than `mq.ErrQueueFull` may have stored the event, so that record's claim is left to lapse and the idempotency key drops the retry's copy. Each row goes through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). +- **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates and encodes each record, and runs the records in windows of up to 256 (`ingestWindow`) through three phases: one dedupe `Reserve` for the window's ids, the publishes in record order (a deduped record under `mq.WithIdempotencyKey`, keyed by `dedupe.IdempotencyKey`), and one `Commit` of the published ids — a window is the unit of a dedupe round trip and of Pebble's commit `fsync`. An id another request holds answers `503` with the lease as `Retry-After`, a store that cannot answer (`dedupe.ErrUnavailable`) `503` with `Retry-After: 5`; a publish that fails at a record commits the ones before it and releases the rest, except that a failure other than `mq.ErrQueueFull` may have stored the event, so that record's claim is left to lapse and the idempotency key drops the retry's copy if it comes within the stream's two-minute duplicate window. Each row goes through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` (or setting it to `null`) can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index a0e8e36ed..5f84c624e 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -60,9 +60,9 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) ## Deduplication: one more fsync per window -With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids are committed to the dedupe store, and on the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes. The retry that follows the lease therefore stores no second copy. That holds only while the lease is shorter than the stream's duplicate window. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The window covers a prompt retry only while the lease is shorter than it. ## Check your storage before you trust it diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index a9bdeb82b..d10be43ee 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index 21f2e0752..3d1dc931a 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -339,7 +339,7 @@ func TestIngest_Dedup_UncertainPublishThenRetryIsOneEvent(t *testing.T) { } return false, nil }) - h.DedupeLease = 300 * time.Millisecond + h.DedupeLease = 2 * time.Second lines := eventLines(t, 3) send := func() *httptest.ResponseRecorder { w := httptest.NewRecorder() @@ -351,13 +351,13 @@ func TestIngest_Dedup_UncertainPublishThenRetryIsOneEvent(t *testing.T) { require.Equal(t, http.StatusInternalServerError, w.Code) w = send() require.Equal(t, http.StatusServiceUnavailable, w.Code, "the uncertain claim is still held") - assert.Equal(t, "1", w.Header().Get("Retry-After")) + assert.Equal(t, "2", w.Header().Get("Retry-After")) var last *httptest.ResponseRecorder require.Eventually(t, func() bool { last = send() return last.Code == http.StatusOK - }, 5*time.Second, 50*time.Millisecond) + }, 10*time.Second, 100*time.Millisecond) assert.Equal(t, 3, decodeBatchResult(t, last).Succeeded, "the lapsed claim is claimed again and republished") assert.Equal(t, 3, count(), "the republished e1 was dropped by the queue") diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index 004938d02..fc7900b0c 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -143,13 +143,15 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // A repeated idempotency key inside the duplicate window is dropped as a // success, so an uncertain publish can be republished safely; a queue made -// before the window was set gets it on its next budget apply. +// with another window gets this one on its next budget apply. func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { e := openEmbedded(t, t.TempDir()) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() + // Explicit rather than the server's default, which happens to match today. + require.Equal(t, EmbeddedDuplicateWindow, ingestStreamConfig(tenant.Default, testBudget).Duplicates) old := ingestStreamConfig(tenant.Default, testBudget) - old.Duplicates = 0 + old.Duplicates = 10 * time.Second _, err := e.js.CreateStream(ctx, old) require.NoError(t, err) require.NoError(t, e.SetMaxBytes(ctx, tenant.Default, testBudget)) From 9824663384c5c2468583be4fb5f427bb3af50b86 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:26:22 -0400 Subject: [PATCH 03/22] docs(ingest): Release gives back definite failures only; tighten wording Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- 5 files changed, 5 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f0dab22eb..195038a74 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index dd84c32e1..ca5f9b5c3 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -272,7 +272,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 413 | `{"error":"request body exceeded 16777216 bytes"}` | Request body over the 16 MiB cap | | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | -| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store cannot answer now (not open, or a remote backend throttled or unreachable); `Retry-After: 5`. Nothing was published, so the retry is safe | +| 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store is not open (for example, it failed to open on a reload); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue remembers the key for two minutes after the first publish, so a retry inside that window stores no second copy (the SDK's, after the 30-second `Retry-After`, lands inside it); a later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e50b269fd..e0812d7d5 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -122,7 +122,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `dedupe/` — Deduplication (Optional) -- **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. +- **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose records were definitely not published (a refused or never-sent publish; one whose outcome is unknown is left to lapse instead). A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 5f84c624e..280bf458d 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The window covers a prompt retry only while the lease is shorter than it. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. ## Check your storage before you trust it diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index d10be43ee..1e898fdca 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -186,7 +186,7 @@ What stays in boot config is only what cannot change under a running process — Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. From 99a5eee4b9ab7dd9f5fd46903581a298e4ace48d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:49:44 -0400 Subject: [PATCH 04/22] docs(ingest): say the Pebble benchmark stubs the queue; finish the rename Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/durability.md | 2 +- internal/api/ingest.go | 4 ++-- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 195038a74..878072338 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s measured). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 280bf458d..5a736b679 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -60,7 +60,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) ## Deduplication: one more fsync per window -With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, that batch took 24 ms windowed against 5.7 s one record at a time. A single-record request still pays one sync for its publish and one for its commit. +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index acf8bec88..6f129e34c 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -505,7 +505,7 @@ func writeMaxBytesError(w http.ResponseWriter, err error, limit int64) bool { // // Evaluated here rather than per record because the condition is a property of // (table, role, policy) and is identical for every record in the request — the -// same reasoning as the !resolved abort in processRecord. Doing it per record +// same reasoning as the !resolved abort in prepareRecord. Doing it per record // would emit one ERROR line per record for a single mis-wired policy, which on // a 16 MiB body of small records is ~1.2M lines. The reject is still returned // per record, so a batch reports each record's own cause: one that SUPPLIES the @@ -518,7 +518,7 @@ func (h *IngestHandler) policyCheckGuard( ) *recordReject { checks, resolved := perms.CheckClauses() if !resolved { - return nil // the !resolved abort in processRecord owns this case + return nil // the !resolved abort in prepareRecord owns this case } // Sorted, and every offender — not the first one a map range happens to From b83e67075a85017e31ae7d31fabd1b1bf939d35e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 03:49:16 -0400 Subject: [PATCH 05/22] feat(dedupe): retention per tenant and table, and an expiry sweep dedupe.retention (required, "0" = forever) and its per-table override are read per record and passed to Commit. Validation refuses a finite retention below the queue's two-minute duplicate window. The embedded Pebble store deletes expired keys and the version-0 keys in an hourly background sweep, counted by wavehouse_dedupe_swept_keys_total. Part of #613. Refs #220. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 3 +- cmd/wavehouse/validate_test.go | 2 +- config.yaml | 2 +- deployments/compose/settings/config.json | 1 + docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 6 +- docs/src/content/docs/durability.md | 4 +- docs/src/content/docs/settings-directory.mdx | 9 +- internal/api/ingest.go | 66 +++++--- internal/api/ingest_retention_test.go | 104 ++++++++++++ internal/api/ingest_test.go | 40 +++-- internal/api/ingest_window_test.go | 4 +- internal/api/settings_test.go | 2 +- internal/app/app_test.go | 14 +- internal/dedupe/embedded.go | 65 ++++++-- internal/dedupe/sweep.go | 151 ++++++++++++++++++ internal/dedupe/sweep_test.go | 158 +++++++++++++++++++ internal/settings/registry_test.go | 4 +- internal/settings/seed/config.json | 1 + internal/settings/settings.go | 18 ++- internal/settings/store.go | 31 +++- internal/settings/store_test.go | 21 ++- internal/settings/validate.go | 30 +++- internal/settings/validate_test.go | 39 ++++- internal/testutil/mocks.go | 13 +- tests/e2e/fixtures/settings/config.json | 1 + 27 files changed, 693 insertions(+), 102 deletions(-) create mode 100644 internal/api/ingest_retention_test.go create mode 100644 internal/dedupe/sweep.go create mode 100644 internal/dedupe/sweep_test.go diff --git a/AGENTS.md b/AGENTS.md index ffbd3c16f..b26db3e3a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, committed ids stored with their expiry and deleted by an hourly background sweep along with the version-0 keys from before the table joined the key, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal; `WithIdempotencyKey` makes a republish inside the queue's duplicate window a no-op), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` @@ -58,7 +58,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. 6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format`, or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. -8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase, one call per phase per window of up to 256 records — `Reserve` → publish (under the id's idempotency key) → `Commit`, or `Release` when the publish definitely failed, while one whose outcome is unknown is left to lapse; a store that cannot answer is a `503` + `Retry-After`; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. +8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant and table; claims are two-phase, one call per phase per window of up to 256 records — `Reserve` → publish (under the id's idempotency key) → `Commit`, or `Release` when the publish definitely failed, while one whose outcome is unknown is left to lapse; a store that cannot answer is a `503` + `Retry-After`; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key and `dedupe.retention` how long a committed id stays a duplicate (`"0"` = forever, else at least the queue's two-minute duplicate window), both overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 878072338..f47f81ba6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,6 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It deletes 1,024 keys per chunk without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed @@ -78,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, deleted by the retention sweep below ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). A table name holding a NUL byte is refused in `dedupe.tables` and `dlq.tables`. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/cmd/wavehouse/validate_test.go b/cmd/wavehouse/validate_test.go index e2e18e6c0..5598cf650 100644 --- a/cmd/wavehouse/validate_test.go +++ b/cmd/wavehouse/validate_test.go @@ -19,7 +19,7 @@ func writeSettingsDir(t *testing.T, policies string) string { "roles.json": `{"roles": ["public"]}`, "policies.json": policies, "pipes.json": `{}`, - "config.json": `{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 10000, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, + "config.json": `{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 10000, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, } for name, content := range files { require.NoError(t, os.WriteFile(filepath.Join(dir, name), []byte(content), 0o600)) diff --git a/config.yaml b/config.yaml index 53a435029..358124171 100644 --- a/config.yaml +++ b/config.yaml @@ -63,7 +63,7 @@ auth: # clickhouse wiring (addr, http_port, http_scheme, database, username, # query_timeout, tls, headers, max_open_conns, max_idle_conns), auth # (jwks_url, role_claim), dedupe (enabled/id_field/ -# require_id + per-table overrides), dlq.enabled (+ per table), +# require_id/retention + per-table overrides), dlq.enabled (+ per table), # query.default_max_rows / timestamp_bucket_seconds, # schema.refresh_interval, stream keepalive_interval / keepalive_buckets / # gap_window_minutes, mq.max_bytes_gb, cors.allowed_origins — and every key diff --git a/deployments/compose/settings/config.json b/deployments/compose/settings/config.json index 030d76cc1..6b33f55c5 100644 --- a/deployments/compose/settings/config.json +++ b/deployments/compose/settings/config.json @@ -26,6 +26,7 @@ "enabled": false, "id_field": "event_id", "require_id": false, + "retention": "0", "tables": {} }, "dlq": { diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e0812d7d5..29b6b2842 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -124,7 +124,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose records were definitely not published (a refused or never-sent publish; one whose outcome is unknown is left to lapse instead). A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. -- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync, each value carrying its expiry (`0` = never), which `Reserve` honors on read. A background sweep (`sweep.go`), started when the instance opens and stopped before it closes, deletes expired keys and the version-0 keys from before the table joined the key ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)): a minute after opening, then hourly, 1,024 keys per chunk, holding a lock `Commit` also takes, so a key re-committed after the sweep read it is never deleted; `wavehouse_dedupe_swept_keys_total{reason}` counts what it deletes. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index bd47e10a8..86347f444 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -169,7 +169,7 @@ WH_SETTINGS_DIR=/etc/wavehouse/settings WaveHouse keeps all embedded state under a single configurable root, `WH_DATA_DIR` (yaml: `data_dir`). Subdirectories are convention, not config: - `/nats` — embedded NATS JetStream. Holds in-flight events between an ingest POST and the ingest worker → ClickHouse flush, plus the `stream.gap_window_minutes` window (settings directory) of history that powers SSE gap-fill across restarts. -- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). +- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). It grows with every id kept: with `dedupe.retention` at `"0"` (forever) nothing is ever removed, so size the volume for it or set a [retention](/settings-directory#deduplication), whose expired ids an hourly sweep deletes. In a Docker / Podman / Kubernetes deployment, **`data_dir` must resolve to a host-backed volume**. The reference compose file `deployments/compose/standalone.yaml` sets `WH_DATA_DIR=/app/data` and binds a `wavehouse-data:/app/data` volume — copy that pattern. The bundled Dockerfiles pre-create `/app/data` and `/app/settings` owned by the nonroot user (UID 65532); the binary creates the `nats/` and `pebble/` subdirectories under `/app/data` itself on first run. @@ -419,7 +419,9 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres ## Upgrading across the dedupe key change -The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread; nothing removes them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep that will). Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys are never read, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. + +The same release adds **`dedupe.retention`, a required key**: every `config.json`, each tenant's folder included, must state it or the directory is refused (at boot) or not adopted (on reload). `"retention": "0"` keeps every id forever, as before; see [Deduplication](/settings-directory#deduplication) for a finite one. ## Upgrading across the v2 ingest envelope diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 5a736b679..a0a62952c 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,9 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. +With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads and deletes 1,024 keys at a time, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. + +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. ## Check your storage before you trust it diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 1e898fdca..126f35dd1 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -123,7 +123,8 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `dedupe.enabled` | `false` | Turn deduplication on; a reload opens or closes this tenant's store — see [Deduplication](#deduplication). | | `dedupe.id_field` | `event_id` | Dedup key field — see [Deduplication](#deduplication). | | `dedupe.require_id` | `false` | Reject rows missing the id field — see [Deduplication](#deduplication). | -| `dedupe.tables.
.{id_field, require_id}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | +| `dedupe.retention` | `"0"` | How long a committed id stays a duplicate, as a duration (`"720h"`); `"0"` keeps it forever — see [Deduplication](#deduplication). | +| `dedupe.tables.
.{id_field, require_id, retention}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | | `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the tenant's dead-letter stream (`DLQ_{tenant}`) (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | | `dlq.tables.
.enabled` | `{}` | Optional per-table override of the switch. | | `query.timestamp_bucket_seconds` | `60` | Bucket (seconds, `>= 0`) that a structured query's relative time range is truncated to, so near-identical queries share a cache entry; `0` disables bucketing. Read per query. | @@ -161,8 +162,9 @@ The tenant tunables. Every key is required (a missing one is a validation error) "enabled": false, "id_field": "event_id", "require_id": false, + "retention": "720h", "tables": { - "clicks": { "id_field": "click_id" } + "clicks": { "id_field": "click_id", "retention": "24h" } } }, "dlq": { @@ -188,7 +190,8 @@ Every dedupe knob lives here — there are no boot-config keys for it. The switc - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. +- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep deletes the expired id from the store: first a minute after the store opens, then hourly, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"` or a bare number. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 6f129e34c..1189547a4 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -50,12 +50,11 @@ type IngestHandler struct { // store, picked off the store the handler already holds (#583 story 7; // dedupe.Stores in production). nil when no dedupe store is wired (tests). Dedup func(store *settings.Store) dedupe.Deduplicator - // DedupeSettings resolves the effective dedupe id_field/require_id for a - // table of the request's tenant ((*settings.Store).DedupeFor in - // production). Called once per record so a settings reload lands at a - // record boundary — one record never mixes two documents' values. Dedup is - // skipped when nil. - DedupeSettings func(store *settings.Store, table string) (enabled bool, idField string, requireID bool) + // DedupeSettings resolves the effective dedupe settings for a table of the + // request's tenant ((*settings.Store).DedupeFor in production). Called + // once per record so a settings reload lands at a record boundary — one + // record never mixes two documents' values. Dedup is skipped when nil. + DedupeSettings func(store *settings.Store, table string) settings.Dedupe // DedupeLease is how long a record's claimed id stays pending while it is // published; 0 means dedupe.DefaultLease. DedupeLease time.Duration @@ -572,8 +571,10 @@ type pendingRecord struct { reject *recordReject // non-nil: the record is bad and is not published payload []byte // the encoded envelope to publish // key is the record's dedupe identity, nil when it is published - // un-deduped; claim is Reserve's answer for it. + // un-deduped; retention is how long its id stays a duplicate once + // committed; claim is Reserve's answer for it. key *dedupe.Key + retention time.Duration claim dedupe.Claim duplicate bool } @@ -701,7 +702,7 @@ func (h *IngestHandler) prepareRecord( // enforces) after the permission checks: check clauses keep pre-#372 semantics. h.validator().CanonicalizeTimestamps(schema, data) - // Optional deduplication. enabled/id_field/require_id resolve per record + // Optional deduplication. The dedupe settings resolve per record // from one snapshot (table override → global; the settings directory // always states them, so no compiled fallback is needed), so a reload // lands at a record boundary. A Deduplicator without a settings source is @@ -709,14 +710,16 @@ func (h *IngestHandler) prepareRecord( // in ingestWindow, once every record of the window is encoded, so nothing // but the publish can fail while the claim is held. if h.Dedup != nil && h.DedupeSettings != nil { - if enabled, idField, requireID := h.DedupeSettings(store, table); enabled { + if dd := h.DedupeSettings(store, table); dd.Enabled { + idField := dd.IDField // An explicit null is as missing as an absent key (#370): fmt.Sprint // would make every null "", one id for every such record. if idVal, ok := data[idField]; ok && idVal != nil { rec.key = &dedupe.Key{Table: table, ID: fmt.Sprint(idVal)} + rec.retention = dd.Retention } else { dedupeMissingIDCounter.Add(ctx, 1, metric.WithAttributes(attribute.String("table", table))) - if requireID { + if dd.RequireID { slog.WarnContext(ctx, "dedupe id_field missing or null; rejecting", "id_field", idField, "table", table) return pendingRecord{reject: &recordReject{ Status: http.StatusBadRequest, @@ -793,7 +796,7 @@ func (h *IngestHandler) ingestWindow(ctx context.Context, store *settings.Store, return h.publishFailed(ctx, dd, topic, recs, i, err) } } - commitClaims(ctx, dd, claimedIn(recs), table) + commitClaims(ctx, dd, recs, table) return nil } @@ -865,7 +868,7 @@ func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, tab // the first copy landed. The records after k were never sent and are released. func (h *IngestHandler) publishFailed(ctx context.Context, dd dedupe.Deduplicator, topic mq.Topic, recs []pendingRecord, k int, err error) *requestAbort { definite := errors.Is(err, mq.ErrQueueFull) - commitClaims(ctx, dd, claimedIn(recs[:k]), topic.Table) + commitClaims(ctx, dd, recs[:k], topic.Table) after := k + 1 if definite { after = k @@ -890,20 +893,33 @@ func claimedIn(recs []pendingRecord) []dedupe.Claim { return out } -// commitClaims makes published records' ids duplicates. A failure does not -// fail the records — they are in the queue — so it is logged and counted, and -// the claims lapse after their lease. -func commitClaims(ctx context.Context, dd dedupe.Deduplicator, claims []dedupe.Claim, table string) { - if len(claims) == 0 { - return +// commitClaims makes the ids of recs' Claimed claims duplicates, one Commit +// per retention — one in practice, unless a reload changed it mid-window. A +// failure does not fail the records — they are in the queue — so it is logged +// and counted, and the claims lapse after their lease. +func commitClaims(ctx context.Context, dd dedupe.Deduplicator, recs []pendingRecord, table string) { + var retentions []time.Duration + byRetention := map[time.Duration][]dedupe.Claim{} + for i := range recs { + if recs[i].claim.Status != dedupe.Claimed { + continue + } + r := recs[i].retention + if _, ok := byRetention[r]; !ok { + retentions = append(retentions, r) + } + byRetention[r] = append(byRetention[r], recs[i].claim) } - // The records are queued whatever the request's context does next. - err := dd.Commit(context.WithoutCancel(ctx), claims, 0) - switch { - case err == nil, errors.Is(err, dedupe.ErrDisabled): - default: - dedupeCommitFailedCounter.Add(ctx, int64(len(claims)), metric.WithAttributes(attribute.String("table", table))) - slog.ErrorContext(ctx, "dedupe commit failed after publish; the ids lapse with their lease", "error", err, "table", table, "records", len(claims)) + for _, r := range retentions { + claims := byRetention[r] + // The records are queued whatever the request's context does next. + err := dd.Commit(context.WithoutCancel(ctx), claims, r) + switch { + case err == nil, errors.Is(err, dedupe.ErrDisabled): + default: + dedupeCommitFailedCounter.Add(ctx, int64(len(claims)), metric.WithAttributes(attribute.String("table", table))) + slog.ErrorContext(ctx, "dedupe commit failed after publish; the ids lapse with their lease", "error", err, "table", table, "records", len(claims)) + } } } diff --git a/internal/api/ingest_retention_test.go b/internal/api/ingest_retention_test.go new file mode 100644 index 000000000..ef65e4b2f --- /dev/null +++ b/internal/api/ingest_retention_test.go @@ -0,0 +1,104 @@ +package api + +import ( + "net/http" + "net/http/httptest" + "os" + "path/filepath" + "strings" + "sync/atomic" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/discovery" + "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil" +) + +// A finite retention must outlast the queue's duplicate window, or an id +// re-sent after it expires is claimed again and then dropped by the queue as +// a copy of the first publish. +func TestIngest_MinDedupeRetentionCoversTheDuplicateWindow(t *testing.T) { + t.Parallel() + assert.GreaterOrEqual(t, settings.MinDedupeRetention, mq.EmbeddedDuplicateWindow) +} + +// dedupeConfig is fullConfig with dedupe switched on and the given dedupe +// block's retention settings. +func dedupeConfig(retention, tables string) string { + return strings.Replace(fullConfig(100), + `"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}`, + `"dedupe": {"enabled": true, "id_field": "event_id", "require_id": false, "retention": "`+retention+`", "tables": `+tables+`}`, 1) +} + +// Each record is committed with its table's retention from the adopted +// settings, and a reload changes it for the next request: the retention is +// read per record, like id_field, not fixed when the store was opened. +func TestIngest_Dedup_CommitsWithTheAdoptedRetention(t *testing.T) { + t.Parallel() + dir := writeSettingsFixture(t, dedupeConfig("720h", `{"users": {"retention": "0"}}`)) + tenants, findings := settings.Open(dir) + require.NotNil(t, tenants, "findings: %v", findings) + store, _ := tenants.For(tenant.Default) + + reg := testutil.NewTestSchemaRegistry(t, []*discovery.TableSchema{ + {Name: "clicks", Columns: []discovery.Column{{Name: "event_id", Type: "String"}}}, + {Name: "users", Columns: []discovery.Column{{Name: "event_id", Type: "String"}}}, + }) + dedup := testutil.NewMockDeduplicator() + h := NewIngestHandler(fixedRegistry(reg), &testutil.MockPublisher{}) + h.Dedup = staticDedup(dedup) + h.DedupeSettings = (*settings.Store).DedupeFor + ingest := func(table, id string) { + t.Helper() + w := httptest.NewRecorder() + req := ingestRequest(t, table, map[string]any{"event_id": id}) + h.Handle(w, req.WithContext(WithStore(req.Context(), store))) + require.Equal(t, http.StatusOK, w.Code, w.Body.String()) + } + + ingest("clicks", "e1") + ingest("users", "e1") + assert.Equal(t, 720*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "e1"})) + assert.Equal(t, time.Duration(0), dedup.Retention(dedupe.Key{Table: "users", ID: "e1"}), "the table keeps ids forever") + + require.NoError(t, os.WriteFile(filepath.Join(dir, settings.FileConfig), []byte(dedupeConfig("24h", `{}`)), 0o600)) + _, adopted := tenants.Reload("test") + require.True(t, adopted) + ingest("clicks", "e2") + ingest("users", "e2") + assert.Equal(t, 24*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "e2"})) + assert.Equal(t, 24*time.Hour, dedup.Retention(dedupe.Key{Table: "users", ID: "e2"}), "the override is gone") + assert.Equal(t, 720*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "e1"}), "ids committed before the change keep theirs") +} + +// A reload that lands mid-window splits the window's commit by retention, so +// every record keeps the retention of the snapshot it was prepared under. +func TestIngest_Dedup_ReloadMidWindowCommitsEachRetention(t *testing.T) { + t.Parallel() + dedup := testutil.NewMockDeduplicator() + h := NewIngestHandler(fixedRegistry(testRegistry(t)), &testutil.MockPublisher{}) + h.Dedup = staticDedup(dedup) + var calls atomic.Int32 + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + if calls.Add(1) <= 2 { + return settings.Dedupe{Enabled: true, IDField: "event_id", Retention: time.Hour} + } + return settings.Dedupe{Enabled: true, IDField: "event_id", Retention: 2 * time.Hour} + } + + w := httptest.NewRecorder() + h.Handle(w, withTenant(ndjsonRequest(t, "clicks", + `{"page": "/", "event_id": "a"}`, `{"page": "/", "event_id": "b"}`, `{"page": "/", "event_id": "c"}`))) + require.Equal(t, http.StatusOK, w.Code, w.Body.String()) + assert.Equal(t, 2, dedup.Commits, "one Commit per retention") + assert.Equal(t, time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "a"})) + assert.Equal(t, time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "b"})) + assert.Equal(t, 2*time.Hour, dedup.Retention(dedupe.Key{Table: "clicks", ID: "c"})) +} diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index 579933339..763ba2838 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -201,7 +201,9 @@ func TestIngest_Dedup_FirstTime(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "evt-1"}) w := httptest.NewRecorder() @@ -217,7 +219,9 @@ func TestIngest_Dedup_Duplicate(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } // First call. req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "dup-1"}) @@ -703,7 +707,9 @@ func TestIngest_DedupIsTheTenants(t *testing.T) { pub := &testutil.MockPublisher{} h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = func(s *settings.Store) dedupe.Deduplicator { return stores.For(s.Tenant()) } - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } ingest := func(id tenant.ID) string { store, ok := tenants.For(id) @@ -726,7 +732,9 @@ func TestIngest_Dedup_MissingIDField(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } // Payload omits event_id and require_id is off: the row skips // dedup and is still published — the warn+counter path, not a rejection (#219). @@ -745,7 +753,9 @@ func TestIngest_Dedup_RequireID_Rejects(t *testing.T) { pub := &testutil.MockPublisher{} h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(testutil.NewMockDeduplicator()) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: true} + } w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"page": "/home"}))) @@ -767,7 +777,9 @@ func TestIngest_NDJSON_RequireID_Rejects(t *testing.T) { pub := &testutil.MockPublisher{} h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(testutil.NewMockDeduplicator()) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: true} + } req := ndjsonRequest(t, "clicks", jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), @@ -1015,7 +1027,9 @@ func TestIngest_NDJSON_Dedup(t *testing.T) { dedup := testutil.NewMockDeduplicator() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } req := ndjsonRequest(t, "clicks", jsonLine(t, map[string]any{"page": "/a", "event_id": "e1"}), @@ -2306,7 +2320,9 @@ func TestIngest_Dedup_DisabledBySettings(t *testing.T) { dedup.Err = errors.New("must not be called while disabled") h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return false, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{IDField: "event_id", RequireID: true} + } w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", tt.body))) @@ -2326,7 +2342,9 @@ func TestIngest_Dedup_DisabledMidReload(t *testing.T) { dedup.Err = dedupe.ErrDisabled h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", true } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: true} + } w := httptest.NewRecorder() h.Handle(w, withTenant(ingestRequest(t, "clicks", map[string]any{"event_id": "e1", "page": "/home"}))) @@ -2744,7 +2762,9 @@ func dedupHandler(t *testing.T, pub *testutil.MockPublisher, dedup dedupe.Dedupl t.Helper() h := NewIngestHandler(fixedRegistry(testRegistry(t)), pub) h.Dedup = staticDedup(dedup) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", requireID } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id", RequireID: requireID} + } return h } diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index 3d1dc931a..32015dc48 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -391,7 +391,9 @@ func pebbleBatchHandler(tb testing.TB, window int) (*IngestHandler, *countingDed counted := &countingDedup{Deduplicator: store} h := NewIngestHandler(fixedRegistry(testRegistry(tb)), &testutil.MockPublisher{}) h.Dedup = staticDedup(counted) - h.DedupeSettings = func(*settings.Store, string) (bool, string, bool) { return true, "event_id", false } + h.DedupeSettings = func(*settings.Store, string) settings.Dedupe { + return settings.Dedupe{Enabled: true, IDField: "event_id"} + } h.window = window return h, counted } diff --git a/internal/api/settings_test.go b/internal/api/settings_test.go index 9e823ceaa..24fa80f12 100644 --- a/internal/api/settings_test.go +++ b/internal/api/settings_test.go @@ -19,7 +19,7 @@ import ( // fullConfig is a complete config.json (every key is required) with the // given query.default_max_rows. func fullConfig(maxRows int) string { - return fmt.Sprintf(`{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": %d, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, maxRows) + return fmt.Sprintf(`{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": %d, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, maxRows) } // writeSettingsFixture materializes a minimal valid settings directory whose diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 1468d0df5..019214c0a 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -208,7 +208,7 @@ func TestNew_DedupeFollowsSettings(t *testing.T) { for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { dir := writeSettings(t, map[string]any{"dedupe": map[string]any{ - "enabled": tt.enabled, "id_field": "event_id", "require_id": false, "tables": map[string]any{}, + "enabled": tt.enabled, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}, }}) cfg := testConfig(t, dir) a := newApp(t, cfg, Options{}) @@ -241,7 +241,7 @@ func TestReload_DrivesTheRegisteredHooks(t *testing.T) { require.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, dir, map[string]any{ - "dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}, + "dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}, "mq": map[string]any{"max_bytes_gb": 2}, }) _, adopted := a.tenants.Reload("test") @@ -405,7 +405,7 @@ func TestNew_NestedWithoutAnOperatorKeyWarnsTheOpsTreeIsClosed(t *testing.T) { // request, so a lost 0 folder is felt at once on the routes that read tenant // 0's list. func TestReload_NestedHooksFollowEachTenant(t *testing.T) { - dedupeOn := map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}} + dedupeOn := map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}} grown := map[string]any{"dedupe": dedupeOn, "mq": map[string]any{"max_bytes_gb": 2}} root := writeNestedSettings(t, map[string]map[string]any{ "0": {"mq": map[string]any{"max_bytes_gb": 1}}, @@ -478,7 +478,7 @@ func TestReload_NestedHooksFollowEachTenant(t *testing.T) { // reopened over the same seen ids when the folder is back. The instance is // open while some tenant's store is, and Close releases it. func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { - dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}} + dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}} root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil, "broken": invalidQuery}) cfg := testConfig(t, root) a := newApp(t, cfg, Options{}) @@ -542,7 +542,7 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { // instance — their ingest answers 500 until a reload or a restart opens it — // while the process, and every tenant with dedupe off, carries on. func TestNew_DedupeOpenFailure(t *testing.T) { - dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}} + dedupeOn := map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}}} // A regular file where the instance's directory should be is what Pebble // refuses to open. block := func(t *testing.T, dataDir string) { @@ -852,7 +852,7 @@ func analystPipe(t *testing.T, dir string) { func TestNew_LateBootFailureReleasesEverything(t *testing.T) { guardGlobals(t) dir := writeSettings(t, map[string]any{"dedupe": map[string]any{ - "enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}, + "enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}, }}) cfg := testConfig(t, dir) natsDir := filepath.Join(cfg.DataDir, "nats") @@ -1458,7 +1458,7 @@ func TestReload_CeilingRefusesAThirdTupleThenOpensIt(t *testing.T) { func TestReload_TenantGoneReleasesItsPoolAndRegistry(t *testing.T) { jwks, _, fetches := jwksServer(t, "acme-1") acmeSettings := authPatch(jwks.URL) - acmeSettings["dedupe"] = map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}} + acmeSettings["dedupe"] = map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "retention": "0", "tables": map[string]any{}} root := writeNestedSettings(t, map[string]map[string]any{"acme": acmeSettings, "globex": nil}) a := newApp(t, testConfig(t, root), Options{}) acme, acmeRegistry, acmeDedup := a.pools.For("acme"), a.discoveries.For("acme"), a.dedup.For("acme") diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index b3f015f0d..5186a7380 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -30,9 +30,16 @@ import ( type Embedded struct { dir string - mu sync.Mutex // guards db and open - db *pebble.DB - open int // tenant stores open over db + mu sync.Mutex // guards db, open and stopSweep + db *pebble.DB + open int // tenant stores open over db + stopSweep func() // stops db's sweep + + // commitMu is read-held by Commit and held by a sweep chunk, so a sweep + // never deletes a key a Commit rewrote after the sweep read it. + commitMu sync.RWMutex + sweepFirst time.Duration + sweepEvery time.Duration pending *pendingSet tokens atomic.Uint64 @@ -45,7 +52,13 @@ type Embedded struct { // NewEmbedded returns the embedded implementation under dataDir. Nothing is // opened until a tenant's store is. func NewEmbedded(dataDir string) *Embedded { - return &Embedded{dir: filepath.Join(dataDir, "pebble"), pending: newPendingSet(), now: time.Now} + return &Embedded{ + dir: filepath.Join(dataDir, "pebble"), + pending: newPendingSet(), + now: time.Now, + sweepFirst: sweepFirstDelay, + sweepEvery: sweepInterval, + } } // Dir is where the instance lives. @@ -76,6 +89,7 @@ func (e *Embedded) acquire(prefix []byte) (Deduplicator, error) { return nil, err } e.db = db + e.stopSweep = e.startSweep(db) } e.open++ return &tenantStore{e: e, db: e.db, prefix: prefix}, nil @@ -90,6 +104,7 @@ func (e *Embedded) release() error { if e.open > 0 { return nil } + e.stopSweep() err := e.db.Close() e.db = nil return err @@ -182,24 +197,36 @@ func (s *tenantStore) reserve(key []byte, k Key, now time.Time, lease time.Durat // committedLive reports whether a stored value is a commit that has not // expired. func committedLive(val []byte, now time.Time) bool { + exp, ok := committedExpiry(val) + return ok && (exp == 0 || now.UnixNano() < exp) +} + +// committedExpired reports whether a stored value is a commit whose +// retention has ended — what the sweep deletes. +func committedExpired(val []byte, now time.Time) bool { + exp, ok := committedExpiry(val) + return ok && exp != 0 && now.UnixNano() >= exp +} + +// committedExpiry reads a commit's expiry (UnixNano, 0 = never); ok is false +// for a value that is not a commit. +func committedExpiry(val []byte) (exp int64, ok bool) { if len(val) != valueLen || val[0] != committedMark { - return false + return 0, false } - exp := int64(binary.BigEndian.Uint64(val[1:])) //nolint:gosec // written from an int64 below - return exp == 0 || now.UnixNano() < exp + return int64(binary.BigEndian.Uint64(val[1:])), true //nolint:gosec // written from an int64 below } // Commit writes every claim in one batch and one fsync, then drops the // pending entries it still owns — in that order, so no Reserve in between // finds the key neither pending nor committed. func (s *tenantStore) Commit(_ context.Context, claims []Claim, retention time.Duration) error { - var exp int64 - if retention > 0 { - exp = s.e.now().Add(retention).UnixNano() - } + s.e.commitMu.RLock() + defer s.e.commitMu.RUnlock() + exp := expiry(s.e.now(), retention) val := make([]byte, valueLen) val[0] = committedMark - binary.BigEndian.PutUint64(val[1:], uint64(exp)) + binary.BigEndian.PutUint64(val[1:], uint64(exp)) //nolint:gosec // expiry is never negative b := s.db.NewBatch() defer func() { _ = b.Close() }() for _, c := range claims { @@ -214,6 +241,20 @@ func (s *tenantStore) Commit(_ context.Context, claims []Claim, retention time.D return nil } +// expiry is the stored expiry of a commit at now kept for retention: 0 for +// none, and the latest representable instant for a retention reaching past +// it, rather than a wrapped-around one in the past. +func expiry(now time.Time, retention time.Duration) int64 { + if retention <= 0 { + return 0 + } + n := now.UnixNano() + if retention > time.Duration(math.MaxInt64-n) { + return math.MaxInt64 + } + return n + int64(retention) +} + // Release drops the pending entries the claims still own. func (s *tenantStore) Release(_ context.Context, claims []Claim) error { s.release(claims) diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go new file mode 100644 index 000000000..e536c5edc --- /dev/null +++ b/internal/dedupe/sweep.go @@ -0,0 +1,151 @@ +package dedupe + +import ( + "bytes" + "context" + "fmt" + "log/slog" + "time" + + "github.com/cockroachdb/pebble" + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/metric" +) + +// The sweep's cadence. Expired keys are already absent to Reserve, so the +// sweep only reclaims space and can run rarely; the first pass comes soon +// after the instance opens so an upgrade's version-0 keys go without waiting +// an hour. +const ( + sweepInterval = time.Hour + sweepFirstDelay = time.Minute + // sweepChunk keys are read and deleted per lock hold, with sweepPause + // between chunks: at most ~100k keys a second, and a Commit never waits + // longer than one chunk. + sweepChunk = 1024 + sweepPause = 10 * time.Millisecond +) + +// Swept-key reasons, the metric's reason attribute. +const ( + sweptExpired = "expired" + sweptVersion0 = "version_0" + sweptAttribute = "reason" +) + +var sweptKeysCounter, _ = otel.Meter("wavehouse-dedupe").Int64Counter( + "wavehouse_dedupe_swept_keys_total", + metric.WithDescription("Keys the embedded dedupe sweep deleted, by reason: expired (retention ended) or version_0 (the layout before ids were keyed by table)"), +) + +// sweepResult is what a sweep deleted. +type sweepResult struct { + Expired, Version0 int +} + +// startSweep runs the sweep over db until the returned stop is called; stop +// waits for a chunk in progress to finish. Callers hold e.mu. +func (e *Embedded) startSweep(db *pebble.DB) (stop func()) { + ctx, cancel := context.WithCancel(context.Background()) + done := make(chan struct{}) + go func() { + defer close(done) + wait := e.sweepFirst + for { + select { + case <-ctx.Done(): + return + case <-time.After(wait): + } + wait = e.sweepEvery + res, err := e.sweep(ctx, db) + switch { + case err != nil && ctx.Err() == nil: + slog.WarnContext(ctx, "dedupe sweep failed; retrying next interval", "error", err, "expired", res.Expired, "version_0", res.Version0) + case res.Expired+res.Version0 > 0: + slog.InfoContext(ctx, "dedupe sweep deleted keys", "expired", res.Expired, "version_0", res.Version0) + } + } + }() + return func() { + cancel() + <-done + } +} + +// sweep makes one pass over the whole instance, deleting keys whose +// retention has ended and version-0 keys (tenant ‖ 0x00 ‖ id, from before +// ids were keyed by table), which nothing reads. It stops early, without +// error, when ctx ends. +func (e *Embedded) sweep(ctx context.Context, db *pebble.DB) (sweepResult, error) { + var res sweepResult + var from []byte + for { + next, err := e.sweepChunk(ctx, db, from, &res) + if err != nil || next == nil { + return res, err + } + from = next + select { + case <-ctx.Done(): + return res, nil + case <-time.After(sweepPause): + } + } +} + +// sweepChunk deletes the sweepable keys among the next sweepChunk keys from +// from, returning where the next chunk starts (nil at the end). It holds +// commitMu, so no Commit lands between reading a key and deleting it: a key +// re-committed after it expired is never deleted with its new value. +func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, res *sweepResult) ([]byte, error) { + e.commitMu.Lock() + defer e.commitMu.Unlock() + now := e.now() + it, err := db.NewIter(&pebble.IterOptions{LowerBound: from}) + if err != nil { + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + b := db.NewBatch() + defer func() { _ = b.Close() }() + var next []byte + var expired, v0 int64 + seen := 0 + for valid := it.First(); valid; valid = it.Next() { + if seen == sweepChunk { + next = bytes.Clone(it.Key()) + break + } + seen++ + k := it.Key() + switch { + case len(k) == 0 || k[0] != keyVersion: + v0++ + case committedExpired(it.Value(), now): + expired++ + default: + continue + } + if err := b.Delete(k, nil); err != nil { + _ = it.Close() + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + } + if err := it.Close(); err != nil { + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + // NoSync: a delete lost to a crash is redone by the next pass. + if err := b.Commit(pebble.NoSync); err != nil { + return nil, fmt.Errorf("dedupe sweep: %w", err) + } + res.Expired += int(expired) + res.Version0 += int(v0) + if expired > 0 { + sweptKeysCounter.Add(ctx, expired, metric.WithAttributes(attribute.String(sweptAttribute, sweptExpired))) + } + if v0 > 0 { + sweptKeysCounter.Add(ctx, v0, metric.WithAttributes(attribute.String(sweptAttribute, sweptVersion0))) + } + return next, nil +} diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go new file mode 100644 index 000000000..44579dcd8 --- /dev/null +++ b/internal/dedupe/sweep_test.go @@ -0,0 +1,158 @@ +package dedupe + +import ( + "context" + "errors" + "fmt" + "math" + "sync/atomic" + "testing" + "time" + + "github.com/cockroachdb/pebble" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// stepClock is a clock a test moves by hand. +type stepClock struct{ ns atomic.Int64 } + +func newStepClock() *stepClock { + c := &stepClock{} + c.ns.Store(time.Now().UnixNano()) + return c +} + +func (c *stepClock) now() time.Time { return time.Unix(0, c.ns.Load()) } +func (c *stepClock) advance(d time.Duration) { c.ns.Add(int64(d)) } +func present(t *testing.T, e *Embedded, key []byte) bool { + t.Helper() + _, closer, err := e.db.Get(key) + if errors.Is(err, pebble.ErrNotFound) { + return false + } + require.NoError(t, err) + _ = closer.Close() + return true +} + +// commitIDs reserves and commits ids in table "events" with retention. +func commitIDs(t *testing.T, m *Managed, retention time.Duration, ids ...string) { + t.Helper() + keys := make([]Key, len(ids)) + for i, id := range ids { + keys[i] = Key{Table: "events", ID: id} + } + claims, err := m.Reserve(context.Background(), keys, DefaultLease) + require.NoError(t, err) + require.NoError(t, m.Commit(context.Background(), claims, retention)) +} + +// A sweep deletes the keys whose retention has ended and every version-0 +// key, across chunk boundaries, and leaves every live key: one kept forever, +// one not yet expired, and one that expired and was committed again. +func TestEmbedded_SweepDeletesExpiredAndVersionZeroKeys(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + clock := newStepClock() + SetClock(e, clock.now) + acme, globex := switchedOn(t, e, "acme"), switchedOn(t, e, "globex") + + // More expired keys than one chunk holds, interleaved with live ones. + var expired, live []string + for i := range 2*sweepChunk + 10 { + expired = append(expired, fmt.Sprintf("x%05d", i)) + live = append(live, fmt.Sprintf("x%05d-live", i)) + } + commitIDs(t, acme, time.Hour, expired...) + commitIDs(t, acme, 3*time.Hour, live...) + commitIDs(t, globex, 0, "forever") + commitIDs(t, globex, time.Hour, "recommitted") + for _, k := range []string{"acme\x00e1", "acme\x00e2", "globex\x00e1"} { + require.NoError(t, e.db.Set([]byte(k), make([]byte, 8), pebble.Sync)) + } + + clock.advance(2 * time.Hour) + commitIDs(t, globex, time.Hour, "recommitted") + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Expired: len(expired), Version0: 3}, res) + + for _, id := range expired { + require.False(t, present(t, e, AppendKey(nil, KeyPrefix("acme"), Key{Table: "events", ID: id})), id) + } + for _, id := range live { + require.True(t, present(t, e, AppendKey(nil, KeyPrefix("acme"), Key{Table: "events", ID: id})), id) + } + assert.True(t, present(t, e, AppendKey(nil, KeyPrefix("globex"), Key{Table: "events", ID: "forever"}))) + assert.True(t, present(t, e, AppendKey(nil, KeyPrefix("globex"), Key{Table: "events", ID: "recommitted"}))) + assert.False(t, present(t, e, []byte("acme\x00e1"))) + assert.False(t, present(t, e, []byte("globex\x00e1"))) + + dup, err := mark(context.Background(), globex, "recommitted") + require.NoError(t, err) + assert.True(t, dup, "the new commit survived the sweep") + res, err = e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{}, res, "a second pass finds nothing") +} + +// A retention is honoured on read before any sweep has run: the key is a +// duplicate until the retention ends and claimable from that instant. +func TestEmbedded_RetentionHonouredOnRead(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + clock := newStepClock() + SetClock(e, clock.now) + m := switchedOn(t, e, "acme") + commitIDs(t, m, time.Hour, "e1") + + clock.advance(time.Hour - time.Nanosecond) + claims, err := m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) + require.NoError(t, err) + assert.Equal(t, Duplicate, claims[0].Status) + + clock.advance(time.Nanosecond) + claims, err = m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) + require.NoError(t, err) + assert.Equal(t, Claimed, claims[0].Status) +} + +// The sweep runs on its own once the instance opens, and stops with it. +func TestEmbedded_SweepRunsWhileOpen(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + e.sweepFirst, e.sweepEvery = time.Millisecond, time.Millisecond + m := e.Tenant("acme") + require.NoError(t, m.Apply(true)) + require.NoError(t, e.db.Set([]byte("acme\x00e1"), make([]byte, 8), pebble.Sync)) + assert.Eventually(t, func() bool { return !present(t, e, []byte("acme\x00e1")) }, 5*time.Second, 5*time.Millisecond) + require.NoError(t, m.Apply(false), "closing waits for the sweep to stop") + assert.False(t, e.Open()) +} + +// A sweep stops between chunks when its context ends. +func TestEmbedded_SweepStopsWhenCancelled(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + switchedOn(t, e, "acme") + b := e.db.NewBatch() + for i := range 3 * sweepChunk { + require.NoError(t, b.Set(fmt.Appendf(nil, "acme\x00%05d", i), nil, nil)) + } + require.NoError(t, b.Commit(pebble.Sync)) + ctx, cancel := context.WithCancel(context.Background()) + cancel() + res, err := e.sweep(ctx, e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Version0: sweepChunk}, res, "one chunk, then the cancellation is seen") +} + +func TestExpiry(t *testing.T) { + t.Parallel() + now := time.Unix(0, 1_000) + assert.Zero(t, expiry(now, 0)) + assert.Zero(t, expiry(now, -time.Second)) + assert.Equal(t, 1_000+int64(time.Hour), expiry(now, time.Hour)) + assert.Equal(t, int64(math.MaxInt64), expiry(now, time.Duration(math.MaxInt64)), "saturates rather than wrapping into the past") +} diff --git a/internal/settings/registry_test.go b/internal/settings/registry_test.go index 06e2a3b9a..3ac222ccb 100644 --- a/internal/settings/registry_test.go +++ b/internal/settings/registry_test.go @@ -113,9 +113,7 @@ func TestRegistry_SurvivesVanishedDirectory(t *testing.T) { assert.False(t, adopted) assert.True(t, HasErrors(findings)) assert.Equal(t, 42, s.DefaultMaxRows()) - _, id, req := s.DedupeFor("clicks") - assert.Equal(t, "event_id", id) - assert.False(t, req) + assert.Equal(t, "event_id", s.DedupeFor("clicks").IDField) } // TestRegistry_AfterAdoptRunsOnlyOnAdoption pins the lifecycle hook contract diff --git a/internal/settings/seed/config.json b/internal/settings/seed/config.json index a8ab41a24..61aa8fee3 100644 --- a/internal/settings/seed/config.json +++ b/internal/settings/seed/config.json @@ -26,6 +26,7 @@ "enabled": false, "id_field": "event_id", "require_id": false, + "retention": "0", "tables": {} }, "dlq": { diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 7da6c3faa..68445543a 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -13,6 +13,8 @@ package settings import ( + "time" + "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" ) @@ -141,14 +143,18 @@ type AuthConfig struct { // (dedupe.Managed, one per tenant, each a share of the one embedded Pebble // instance), so the whole block is tenant-owned. // -// id_field and require_id are required here and optional per table: a table -// override inherits whichever field it doesn't name. An empty, +// id_field, require_id and retention are required here and optional per +// table: a table override inherits whichever field it doesn't name. An empty, // whitespace-only, or whitespace-padded id_field is rejected at every level, // so the effective id_field can never be empty or silently unmatchable. type DedupeConfig struct { Enabled *bool `json:"enabled"` IDField *string `json:"id_field"` RequireID *bool `json:"require_id"` + // Retention is how long a committed id stays a duplicate, as a Go + // duration ("720h"); "0" keeps it forever. A change applies to ids + // committed after it. + Retention *string `json:"retention"` // Tables holds per-table overrides keyed by ClickHouse table name (#222). // Names are format-checked only — existence is schema discovery's runtime // concern, same as policies.json table keys. @@ -160,8 +166,16 @@ type DedupeConfig struct { type TableDedupe struct { IDField *string `json:"id_field,omitempty"` RequireID *bool `json:"require_id,omitempty"` + Retention *string `json:"retention,omitempty"` } +// MinDedupeRetention is the shortest finite dedupe retention: the embedded +// queue's duplicate window (mq.EmbeddedDuplicateWindow). A record is +// published under an idempotency key derived from its id, so an id re-sent +// after a shorter retention but inside the window is claimed again and then +// dropped by the queue as a copy, while the client is told it was accepted. +const MinDedupeRetention = 2 * time.Minute + // DLQConfig gates the Dead Letter Queue: whether a row that still fails // after the row-by-row isolation retry is parked on the tenant's dead-letter // queue (and its original acked) or left unacked to be redelivered diff --git a/internal/settings/store.go b/internal/settings/store.go index f68a2fbf3..b3615efef 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -76,23 +76,38 @@ func (s *Store) DedupeEnabled() bool { return *s.doc().Config.Dedupe.Enabled } +// Dedupe is a table's effective dedupe settings. +type Dedupe struct { + Enabled bool + IDField string + RequireID bool + // Retention is how long a committed id stays a duplicate; 0 is forever. + Retention time.Duration +} + // DedupeFor resolves the effective dedupe settings for a table: the switch, // then the table override for each field it names, the global value -// otherwise. All three resolve from one snapshot load, so a reload can never -// hand a record the id_field of one document and the require_id (or enabled) -// of another. -func (s *Store) DedupeFor(table string) (enabled bool, idField string, requireID bool) { +// otherwise. Every field resolves from one snapshot load, so a reload can +// never hand a record the id_field of one document and the require_id, +// retention or switch of another. +func (s *Store) DedupeFor(table string) Dedupe { d := s.doc().Config.Dedupe - enabled, idField, requireID = *d.Enabled, *d.IDField, *d.RequireID + out := Dedupe{Enabled: *d.Enabled, IDField: *d.IDField, RequireID: *d.RequireID} + retention := *d.Retention if td, ok := d.Tables[table]; ok { if td.IDField != nil { - idField = *td.IDField + out.IDField = *td.IDField } if td.RequireID != nil { - requireID = *td.RequireID + out.RequireID = *td.RequireID + } + if td.Retention != nil { + retention = *td.Retention } } - return enabled, idField, requireID + // Validate has parsed it already. + out.Retention, _ = time.ParseDuration(retention) + return out } // ClickHouse is the adopted connection wiring, resolved as one value from diff --git a/internal/settings/store_test.go b/internal/settings/store_test.go index abcc6c902..e96898247 100644 --- a/internal/settings/store_test.go +++ b/internal/settings/store_test.go @@ -46,23 +46,22 @@ func TestStore_Tenant(t *testing.T) { func TestStore_DedupeFor_Cascade(t *testing.T) { t.Parallel() s := newLoadedStore(t, map[string]string{ - FileConfig: configJSON(`{"dedupe": {"require_id": true, "tables": {"clicks": {"id_field": "click_id"}, "views": {"require_id": false}}}}`), + FileConfig: configJSON(`{"dedupe": {"require_id": true, "retention": "720h", "tables": {"clicks": {"id_field": "click_id"}, "views": {"require_id": false, "retention": "24h"}, "audit": {"retention": "0"}}}}`), }) tests := []struct { - name, table, wantID string - wantRequire bool + name, table string + want Dedupe }{ - {name: "table overrides id_field, inherits require_id", table: "clicks", wantID: "click_id", wantRequire: true}, - {name: "table overrides require_id, inherits id_field", table: "views", wantID: "event_id", wantRequire: false}, - {name: "unlisted table gets globals", table: "other", wantID: "event_id", wantRequire: true}, + {name: "table overrides id_field, inherits the rest", table: "clicks", want: Dedupe{IDField: "click_id", RequireID: true, Retention: 720 * time.Hour}}, + {name: "table overrides require_id and retention, inherits id_field", table: "views", want: Dedupe{IDField: "event_id", Retention: 24 * time.Hour}}, + {name: "table keeps ids forever under a finite tenant retention", table: "audit", want: Dedupe{IDField: "event_id", RequireID: true}}, + {name: "unlisted table gets globals", table: "other", want: Dedupe{IDField: "event_id", RequireID: true, Retention: 720 * time.Hour}}, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { t.Parallel() - _, id, req := s.DedupeFor(tt.table) - assert.Equal(t, tt.wantID, id) - assert.Equal(t, tt.wantRequire, req) + assert.Equal(t, tt.want, s.DedupeFor(tt.table)) }) } } @@ -83,9 +82,7 @@ func TestStore_SeedIsValid(t *testing.T) { // decision (deployments/compose/settings ships the opt-in trial one). assert.Len(t, findings, 1, "findings: %s", findingStrings(findings)) assert.Contains(t, findingStrings(findings), "no policy") - _, id, req := s.DedupeFor("anything") - assert.Equal(t, "event_id", id) - assert.False(t, req) + assert.Equal(t, Dedupe{IDField: "event_id"}, s.DedupeFor("anything"), "retention 0: ids kept forever, as before retention existed") assert.Equal(t, ClickHouse{Addr: "localhost:9000", HTTPPort: 8123, HTTPScheme: "http", Database: "default", Username: "default", QueryTimeout: 30 * time.Second, Headers: map[string]string{}, MaxOpenConns: 10, MaxIdleConns: 5}, s.ClickHouse()) assert.Equal(t, Auth{JWKSURL: "", RoleClaim: "role"}, s.Auth()) assert.True(t, s.DLQFor("anything")) diff --git a/internal/settings/validate.go b/internal/settings/validate.go index 08575c953..76352e199 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -12,6 +12,7 @@ import ( "path/filepath" "slices" "strings" + "time" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" @@ -450,6 +451,26 @@ func (v *validator) checkIDField(path string, val *string) { } } +// checkRetention rejects a dedupe retention that is not a duration, is +// negative, or is finite but shorter than MinDedupeRetention. The short one +// is refused rather than raised to the minimum, so the file never means +// something other than what it says. nil is the caller's concern, as for +// id_field. +func (v *validator) checkRetention(path string, val *string) { + if val == nil { + return + } + d, err := time.ParseDuration(*val) + switch { + case err != nil: + v.errorf(FileConfig, path, "must be a duration such as \"720h\", or \"0\" to keep ids forever, got %q", *val) + case d < 0: + v.errorf(FileConfig, path, "must not be negative, got %q", *val) + case d > 0 && d < MinDedupeRetention: + v.errorf(FileConfig, path, "%q is shorter than the ingest queue's %s duplicate window: an id re-sent after it expires but inside the window would be dropped by the queue while the client is told it was accepted — use at least %q, or \"0\" to keep ids forever", *val, MinDedupeRetention, MinDedupeRetention.String()) + } +} + // checkTableName rejects a per-table override key that could never match a // table: empty, carrying surrounding whitespace, or holding NUL. Shared by // the dedupe and dlq override maps. @@ -673,15 +694,20 @@ func (v *validator) parseConfig(data []byte) TenantConfig { if d.RequireID == nil { v.required("dedupe.require_id") } + if d.Retention == nil { + v.required("dedupe.retention") + } v.checkIDField("dedupe.id_field", d.IDField) + v.checkRetention("dedupe.retention", d.Retention) // Sorted iteration keeps finding order deterministic across runs. for _, table := range slices.Sorted(maps.Keys(d.Tables)) { td := d.Tables[table] path := "dedupe.tables." + table v.checkTableName("dedupe.tables", table) v.checkIDField(path+".id_field", td.IDField) - if td.IDField == nil && td.RequireID == nil { - v.warnf(FileConfig, path, "override sets nothing — remove it, or set id_field or require_id") + v.checkRetention(path+".retention", td.Retention) + if td.IDField == nil && td.RequireID == nil && td.Retention == nil { + v.warnf(FileConfig, path, "override sets nothing — remove it, or set id_field, require_id or retention") } } } diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index 4dae9c9d4..a6d7c5f73 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -269,13 +269,13 @@ func TestValidate_ContentRules(t *testing.T) { {"negative max rows", FileConfig, `{"query": {"default_max_rows": -1}}`, "must be >= 1"}, {"zero max rows", FileConfig, `{"query": {"default_max_rows": 0}}`, "must be >= 1"}, {"missing dedupe block", FileConfig, `{"dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "dedupe: required"}, - {"missing dlq block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "dlq: required"}, + {"missing dlq block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "dlq: required"}, {"missing dlq.enabled", FileConfig, `{"dlq": {"tables": {}}}`, "dlq.enabled: required"}, {"empty dlq override table name", FileConfig, `{"dlq": {"tables": {"": {"enabled": false}}}}`, "table name must not be empty"}, {"dlq override table whitespace", FileConfig, `{"dlq": {"tables": {"clicks ": {"enabled": false}}}}`, "surrounding whitespace"}, {"missing query.timestamp_bucket_seconds", FileConfig, `{"query": {"default_max_rows": 1}}`, "query.timestamp_bucket_seconds: required"}, {"negative timestamp bucket", FileConfig, `{"query": {"timestamp_bucket_seconds": -1}}`, "must be >= 0"}, - {"missing stream block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "cors": {"allowed_origins": []}}`, "stream: required"}, + {"missing stream block", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "cors": {"allowed_origins": []}}`, "stream: required"}, {"missing stream.keepalive_interval", FileConfig, `{"stream": {"keepalive_buckets": 3, "gap_window_minutes": 15}}`, "stream.keepalive_interval: required"}, {"missing stream.keepalive_buckets", FileConfig, `{"stream": {"keepalive_interval": 30, "gap_window_minutes": 15}}`, "stream.keepalive_buckets: required"}, {"missing stream.gap_window_minutes", FileConfig, `{"stream": {"keepalive_interval": 30, "keepalive_buckets": 3}}`, "stream.gap_window_minutes: required"}, @@ -284,14 +284,21 @@ func TestValidate_ContentRules(t *testing.T) { {"negative gap window", FileConfig, `{"stream": {"gap_window_minutes": -1}}`, "stream.gap_window_minutes: must be >= 0"}, {"keepalive as a duration string", FileConfig, `{"stream": {"keepalive_interval": "30s"}}`, "keepalive_interval"}, {"missing dedupe.require_id", FileConfig, `{"dedupe": {"id_field": "event_id"}}`, "dedupe.require_id: required"}, - {"missing dedupe.enabled", FileConfig, `{"dedupe": {"id_field": "event_id", "require_id": false}}`, "dedupe.enabled: required"}, + {"missing dedupe.enabled", FileConfig, `{"dedupe": {"id_field": "event_id", "require_id": false, "retention": "0"}}`, "dedupe.enabled: required"}, + {"missing dedupe.retention", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}}`, "dedupe.retention: required"}, + {"dedupe.retention not a duration", FileConfig, configJSON(`{"dedupe": {"retention": "30d"}}`), `dedupe.retention: must be a duration such as "720h"`}, + {"dedupe.retention a number", FileConfig, configJSON(`{"dedupe": {"retention": 3600}}`), "retention"}, + {"dedupe.retention negative", FileConfig, configJSON(`{"dedupe": {"retention": "-1h"}}`), "dedupe.retention: must not be negative"}, + {"dedupe.retention under the duplicate window", FileConfig, configJSON(`{"dedupe": {"retention": "1m59s"}}`), `dedupe.retention: "1m59s" is shorter than the ingest queue's 2m0s duplicate window`}, + {"override retention under the duplicate window", FileConfig, configJSON(`{"dedupe": {"tables": {"clicks": {"retention": "30s"}}}}`), "dedupe.tables.clicks.retention: \"30s\" is shorter"}, + {"override retention not a duration", FileConfig, configJSON(`{"dedupe": {"tables": {"clicks": {"retention": "forever"}}}}`), "dedupe.tables.clicks.retention: must be a duration"}, {"missing query.default_max_rows", FileConfig, `{"query": {}}`, "query.default_max_rows: required"}, {"missing schema.refresh_interval", FileConfig, `{"schema": {}}`, "schema.refresh_interval: required"}, {"missing cors.allowed_origins", FileConfig, `{"cors": {}}`, "cors.allowed_origins: required"}, {"empty config document", FileConfig, `{}`, "cors: required"}, {"otel is boot config", FileConfig, `{"otel": {"enabled": true}}`, "unknown field"}, - {"missing clickhouse block", FileConfig, `{"auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "clickhouse: required"}, - {"missing auth block", FileConfig, `{"clickhouse": {"addr": "h:9000", "http_port": 8123, "http_scheme": "http", "database": "d", "username": "u", "query_timeout": 1}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "auth: required"}, + {"missing clickhouse block", FileConfig, `{"auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "clickhouse: required"}, + {"missing auth block", FileConfig, `{"clickhouse": {"addr": "h:9000", "http_port": 8123, "http_scheme": "http", "database": "d", "username": "u", "query_timeout": 1}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": 1, "timestamp_bucket_seconds": 0}, "schema": {"refresh_interval": 1}, "stream": {"keepalive_interval": 1, "keepalive_buckets": 1, "gap_window_minutes": 0}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": []}}`, "auth: required"}, {"missing clickhouse.addr", FileConfig, `{"clickhouse": {"http_port": 8123, "http_scheme": "http", "database": "d", "username": "u", "query_timeout": 1}}`, "clickhouse.addr: required"}, {"clickhouse.addr without port", FileConfig, `{"clickhouse": {"addr": "localhost"}}`, "must be host:port"}, {"clickhouse.http_port out of range", FileConfig, `{"clickhouse": {"http_port": 70000}}`, "clickhouse.http_port: must be in 1-65535"}, @@ -344,6 +351,28 @@ func TestValidate_ContentRules(t *testing.T) { } } +// A finite retention at or above the duplicate window is accepted, "0" (or +// any zero duration) is forever, and a table may keep ids longer or shorter +// than the tenant, or forever under a finite tenant retention. +func TestValidate_DedupeRetentionAccepted(t *testing.T) { + t.Parallel() + for _, patch := range []string{ + `{"dedupe": {"retention": "0"}}`, + `{"dedupe": {"retention": "0s"}}`, + `{"dedupe": {"retention": "2m"}}`, + `{"dedupe": {"retention": "720h", "tables": {"clicks": {"retention": "24h"}, "views": {"retention": "0"}}}}`, + } { + t.Run(patch, func(t *testing.T) { + t.Parallel() + files := validFiles() + files[FileConfig] = configJSON(patch) + doc, findings := ValidateDir(writeDir(t, files)) + require.NotNil(t, doc, "findings: %s", findingStrings(findings)) + assert.False(t, HasErrors(findings)) + }) + } +} + // TestValidate_ClickHouseTLSPathsAreNotOpened pins that the tls block is // checked for shape only: Validate is pure and also runs on the control // plane, so paths that exist nowhere still validate, and the values reach diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index 624fe0efd..b82974e7e 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -116,6 +116,7 @@ func (m *MockSubscriber) Close() error { return nil } type MockDeduplicator struct { mu sync.Mutex committed map[dedupe.Key]bool + retention map[dedupe.Key]time.Duration // each commit's retention pending map[dedupe.Key]string tokens int // Err, if set, fails Reserve — after ErrAfter calls have succeeded; @@ -132,7 +133,7 @@ type MockDeduplicator struct { var _ dedupe.Deduplicator = (*MockDeduplicator)(nil) func NewMockDeduplicator() *MockDeduplicator { - return &MockDeduplicator{committed: map[dedupe.Key]bool{}, pending: map[dedupe.Key]string{}} + return &MockDeduplicator{committed: map[dedupe.Key]bool{}, retention: map[dedupe.Key]time.Duration{}, pending: map[dedupe.Key]string{}} } // Reserve answers Duplicate for a key repeated in one call, as Managed does. @@ -163,7 +164,7 @@ func (m *MockDeduplicator) Reserve(_ context.Context, keys []dedupe.Key, _ time. return claims, nil } -func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ time.Duration) error { +func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, retention time.Duration) error { m.mu.Lock() defer m.mu.Unlock() m.Commits++ @@ -173,6 +174,7 @@ func (m *MockDeduplicator) Commit(_ context.Context, claims []dedupe.Claim, _ ti for _, c := range claims { if c.Status == dedupe.Claimed { m.committed[c.Key] = true + m.retention[c.Key] = retention delete(m.pending, c.Key) } } @@ -209,6 +211,13 @@ func (m *MockDeduplicator) Committed(k dedupe.Key) bool { return m.committed[k] } +// Retention is the retention k was last committed with. +func (m *MockDeduplicator) Retention(k dedupe.Key) time.Duration { + m.mu.Lock() + defer m.mu.Unlock() + return m.retention[k] +} + // Pending reports whether k is claimed and neither committed nor released. func (m *MockDeduplicator) Pending(k dedupe.Key) bool { m.mu.Lock() diff --git a/tests/e2e/fixtures/settings/config.json b/tests/e2e/fixtures/settings/config.json index a0d15cfce..a8d7c7367 100644 --- a/tests/e2e/fixtures/settings/config.json +++ b/tests/e2e/fixtures/settings/config.json @@ -26,6 +26,7 @@ "enabled": true, "id_field": "event_id", "require_id": false, + "retention": "0", "tables": {} }, "dlq": { From 6c74b65969b26e6cc5bc94ac6957f001ffbed8cb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:16:23 -0400 Subject: [PATCH 06/22] test(dedupe): pin that a commit mid-sweep-chunk survives; docs wording Adds a sweep hook so a Commit can race into the gap between a chunk's read and delete; the test fails (3/3) with commitMu removed. Docs: the sweep follows the shared instance, reads (not deletes) 1,024 keys per chunk, and "0" is the one unitless retention. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/durability.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/dedupe/embedded.go | 3 ++ internal/dedupe/sweep.go | 3 ++ internal/dedupe/sweep_test.go | 35 ++++++++++++++++++++ 6 files changed, 44 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f47f81ba6..03affb5f8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,7 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. -- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It deletes 1,024 keys per chunk without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index a0a62952c..2ad53395b 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads and deletes 1,024 keys at a time, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. +With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads 1,024 keys at a time, deleting the expired ones, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 126f35dd1..b9d301626 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -190,7 +190,7 @@ Every dedupe knob lives here — there are no boot-config keys for it. The switc - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep deletes the expired id from the store: first a minute after the store opens, then hourly, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"` or a bare number. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. - `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 5186a7380..1a620412d 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -47,6 +47,9 @@ type Embedded struct { // readHook, when set, runs before each Pebble read in Reserve; a test // makes it fail to exercise Reserve's all-or-nothing error path. readHook func() error + // sweepHook, when set, runs in a sweep chunk between reading its keys and + // deleting them; a test races a Commit into that gap. + sweepHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index e536c5edc..396d64556 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -135,6 +135,9 @@ func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, r if err := it.Close(); err != nil { return nil, fmt.Errorf("dedupe sweep: %w", err) } + if e.sweepHook != nil { + e.sweepHook() + } // NoSync: a delete lost to a crash is redone by the next pass. if err := b.Commit(pebble.NoSync); err != nil { return nil, fmt.Errorf("dedupe sweep: %w", err) diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index 44579dcd8..0589c1030 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -97,6 +97,41 @@ func TestEmbedded_SweepDeletesExpiredAndVersionZeroKeys(t *testing.T) { assert.Equal(t, sweepResult{}, res, "a second pass finds nothing") } +// A Commit that arrives while a sweep chunk has read an expired key but not +// yet deleted it waits for the chunk, so the new commit is never deleted with +// the old value. Without the lock the Commit lands in the gap and the sweep +// then deletes it; the wait below only ever lets that pass, never fail. +func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + clock := newStepClock() + SetClock(e, clock.now) + m := switchedOn(t, e, "acme") + commitIDs(t, m, time.Hour, "e1") + clock.advance(2 * time.Hour) + + claims, err := m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) + require.NoError(t, err) + require.Equal(t, Claimed, claims[0].Status, "expired: claimable again") + done := make(chan error, 1) + e.sweepHook = func() { + go func() { done <- m.Commit(context.Background(), claims, time.Hour) }() + select { + case <-done: + done <- nil + case <-time.After(50 * time.Millisecond): + } + } + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Expired: 1}, res) + require.NoError(t, <-done) + + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.True(t, dup, "the commit made mid-chunk survived the sweep") +} + // A retention is honoured on read before any sweep has run: the key is a // duplicate until the retention ends and claimable from that instant. func TestEmbedded_RetentionHonouredOnRead(t *testing.T) { From fc4085b1a6f4c20b85a0369b3e4bd6cc51617b37 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 11:25:00 -0400 Subject: [PATCH 07/22] docs(durability): describe the benchmark machine neutrally Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/durability.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 5a736b679..dc6ff7f32 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -60,7 +60,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) ## Deduplication: one more fsync per window -With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on an Apple M4 Pro under a load average near 30, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. +With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. From 7b9595e35cd9a5d2aca3ab5ef50cd7527839f5ec Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:25:03 -0400 Subject: [PATCH 08/22] feat(settings): a missing dedupe.retention keeps ids forever MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit dedupe.retention was required in every config.json. It is now optional: missing means "0" (forever) at the tenant level, and a table override without it inherits the tenant's, as for the other override fields. A present value is validated as before — unparseable, negative, or finite below the queue's duplicate window is refused — so an existing directory needs no change to upgrade. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 8 +++---- internal/settings/settings.go | 11 +++++---- internal/settings/store.go | 5 +++- internal/settings/store_test.go | 17 +++++++++++++ internal/settings/validate.go | 7 ++---- internal/settings/validate_test.go | 25 +++++++++++++++++++- 8 files changed, 59 insertions(+), 18 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 03affb5f8..ffa7cb6fb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,7 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. -- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains **`dedupe.retention`, a required key**, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. **Every existing `config.json` must add it**; `"retention": "0"` changes nothing. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains an optional **`dedupe.retention`** key, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. A `config.json` without the key keeps ids forever, and a table override without one inherits the tenant's, so an existing directory needs no change. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 86347f444..ca23f0fe6 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -421,7 +421,7 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys are never read, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. -The same release adds **`dedupe.retention`, a required key**: every `config.json`, each tenant's folder included, must state it or the directory is refused (at boot) or not adopted (on reload). `"retention": "0"` keeps every id forever, as before; see [Deduplication](/settings-directory#deduplication) for a finite one. +The same release adds an optional **`dedupe.retention`** key. No upgrade step is needed: a `config.json` without it keeps every id forever, as before. See [Deduplication](/settings-directory#deduplication) for a finite one. ## Upgrading across the v2 ingest envelope diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index b9d301626..55a8e0d9c 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -99,7 +99,7 @@ Pipes are read per request, so a reload changes what the next `GET /v1/pipes/{na ## `config.json` keys -The tenant tunables. Every key is required (a missing one is a validation error) except the per-table overrides; the "Seed" column is what `wavehouse bootstrap` writes: +The tenant tunables. Every key is required (a missing one is a validation error) except `dedupe.retention` and the per-table overrides; the "Seed" column is what `wavehouse bootstrap` writes: | Key | Seed | Description | | --- | ---- | ----------- | @@ -123,7 +123,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `dedupe.enabled` | `false` | Turn deduplication on; a reload opens or closes this tenant's store — see [Deduplication](#deduplication). | | `dedupe.id_field` | `event_id` | Dedup key field — see [Deduplication](#deduplication). | | `dedupe.require_id` | `false` | Reject rows missing the id field — see [Deduplication](#deduplication). | -| `dedupe.retention` | `"0"` | How long a committed id stays a duplicate, as a duration (`"720h"`); `"0"` keeps it forever — see [Deduplication](#deduplication). | +| `dedupe.retention` | `"0"` | How long a committed id stays a duplicate, as a duration (`"720h"`); `"0"`, or leaving the key out, keeps it forever — see [Deduplication](#deduplication). | | `dedupe.tables.
.{id_field, require_id, retention}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | | `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the tenant's dead-letter stream (`DLQ_{tenant}`) (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | | `dlq.tables.
.enabled` | `{}` | Optional per-table override of the switch. | @@ -190,8 +190,8 @@ Every dedupe knob lives here — there are no boot-config keys for it. The switc - `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`503 dedupe store unavailable`, `Retry-After: 5`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `503 dedupe store unavailable` (`Retry-After: 5`) until a reload opens it — while the tenants with dedupe off carry on. - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails (counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero), the record is still answered `ok` and the id lapses with its 30-second lease: a retry of it before then answers in-flight, one inside the ingest queue's two-minute duplicate window is dropped there by its idempotency key, and one after that is stored again. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. -- `dedupe.retention` (seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. -- `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. +- `dedupe.retention` (optional; seed default `"0"`) — how long a committed id stays a duplicate, as a Go duration string: `"24h"`, `"720h"` (30 days), `"90m"`. There is no day unit. `"0"` keeps every id forever, which was the only behavior before this key existed, and a `config.json` without the key means the same. Once an id's retention has ended, the next record carrying it is published as new, and a background sweep over the shared Pebble instance deletes the expired id: first about a minute after the instance opens (when the first tenant switches dedupe on), then hourly while any tenant keeps it on, counted by `wavehouse_dedupe_swept_keys_total{reason="expired"}`. A finite retention must be at least `"2m"`, the ingest queue's duplicate window: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and then dropped by the queue as a copy, while the client was told it was accepted. A retention below that is refused, not raised to the minimum; so are a negative value and anything that is not a duration, such as `"30d"`, a number with no unit (`"300"` needs one: `"300s"`; `"0"` is the one exception), or a JSON number rather than a string. Hot-reloadable: a change applies to ids committed after the reload, and an id already committed keeps the expiry it was stored with. +- `dedupe.tables.
.{id_field, require_id, retention}` — per-table overrides; each entry overrides only the fields it names and inherits the rest, so a table with no `retention` keeps the tenant's (forever when the tenant sets none). A table can keep ids for a shorter time than its tenant, or for longer, or forever (`"retention": "0"`) under a finite tenant retention. ## ClickHouse diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 68445543a..eeb03f358 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -143,8 +143,9 @@ type AuthConfig struct { // (dedupe.Managed, one per tenant, each a share of the one embedded Pebble // instance), so the whole block is tenant-owned. // -// id_field, require_id and retention are required here and optional per -// table: a table override inherits whichever field it doesn't name. An empty, +// id_field and require_id are required here; retention is optional, and +// missing means "0" (forever). Every field is optional per table: a table +// override inherits whichever field it doesn't name. An empty, // whitespace-only, or whitespace-padded id_field is rejected at every level, // so the effective id_field can never be empty or silently unmatchable. type DedupeConfig struct { @@ -152,9 +153,9 @@ type DedupeConfig struct { IDField *string `json:"id_field"` RequireID *bool `json:"require_id"` // Retention is how long a committed id stays a duplicate, as a Go - // duration ("720h"); "0" keeps it forever. A change applies to ids - // committed after it. - Retention *string `json:"retention"` + // duration ("720h"); "0", or leaving it out, keeps it forever. A change + // applies to ids committed after it. + Retention *string `json:"retention,omitempty"` // Tables holds per-table overrides keyed by ClickHouse table name (#222). // Names are format-checked only — existence is schema discovery's runtime // concern, same as policies.json table keys. diff --git a/internal/settings/store.go b/internal/settings/store.go index b3615efef..d707451d3 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -93,7 +93,10 @@ type Dedupe struct { func (s *Store) DedupeFor(table string) Dedupe { d := s.doc().Config.Dedupe out := Dedupe{Enabled: *d.Enabled, IDField: *d.IDField, RequireID: *d.RequireID} - retention := *d.Retention + retention := "0" + if d.Retention != nil { + retention = *d.Retention + } if td, ok := d.Tables[table]; ok { if td.IDField != nil { out.IDField = *td.IDField diff --git a/internal/settings/store_test.go b/internal/settings/store_test.go index e96898247..3487ba505 100644 --- a/internal/settings/store_test.go +++ b/internal/settings/store_test.go @@ -1,6 +1,7 @@ package settings import ( + "encoding/json" "os" "path/filepath" "testing" @@ -66,6 +67,22 @@ func TestStore_DedupeFor_Cascade(t *testing.T) { } } +// A config.json without dedupe.retention keeps ids forever, and its table +// overrides inherit that or set their own. +func TestStore_DedupeFor_RetentionMissing(t *testing.T) { + t.Parallel() + var doc map[string]map[string]any + require.NoError(t, json.Unmarshal([]byte(configJSON(`{"dedupe": {"tables": {"clicks": {"id_field": "click_id"}, "views": {"retention": "24h"}}}}`)), &doc)) + delete(doc["dedupe"], "retention") + body, err := json.Marshal(doc) + require.NoError(t, err) + s := newLoadedStore(t, map[string]string{FileConfig: string(body)}) + + assert.Equal(t, Dedupe{IDField: "event_id"}, s.DedupeFor("other"), "forever") + assert.Equal(t, Dedupe{IDField: "click_id"}, s.DedupeFor("clicks"), "inherits forever") + assert.Equal(t, Dedupe{IDField: "event_id", Retention: 24 * time.Hour}, s.DedupeFor("views")) +} + // TestStore_SeedIsValid pins that the shipped starter directory passes its // own gate: `wavehouse bootstrap` must never write something // `wavehouse validate` rejects, and the defaults are readable back. diff --git a/internal/settings/validate.go b/internal/settings/validate.go index 76352e199..52fd7637f 100644 --- a/internal/settings/validate.go +++ b/internal/settings/validate.go @@ -454,8 +454,8 @@ func (v *validator) checkIDField(path string, val *string) { // checkRetention rejects a dedupe retention that is not a duration, is // negative, or is finite but shorter than MinDedupeRetention. The short one // is refused rather than raised to the minimum, so the file never means -// something other than what it says. nil is the caller's concern, as for -// id_field. +// something other than what it says. nil is valid: forever at the tenant +// level, inherited at the table level. func (v *validator) checkRetention(path string, val *string) { if val == nil { return @@ -694,9 +694,6 @@ func (v *validator) parseConfig(data []byte) TenantConfig { if d.RequireID == nil { v.required("dedupe.require_id") } - if d.Retention == nil { - v.required("dedupe.retention") - } v.checkIDField("dedupe.id_field", d.IDField) v.checkRetention("dedupe.retention", d.Retention) // Sorted iteration keeps finding order deterministic across runs. diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index a6d7c5f73..04894acc6 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -285,7 +285,6 @@ func TestValidate_ContentRules(t *testing.T) { {"keepalive as a duration string", FileConfig, `{"stream": {"keepalive_interval": "30s"}}`, "keepalive_interval"}, {"missing dedupe.require_id", FileConfig, `{"dedupe": {"id_field": "event_id"}}`, "dedupe.require_id: required"}, {"missing dedupe.enabled", FileConfig, `{"dedupe": {"id_field": "event_id", "require_id": false, "retention": "0"}}`, "dedupe.enabled: required"}, - {"missing dedupe.retention", FileConfig, `{"dedupe": {"enabled": false, "id_field": "event_id", "require_id": false}}`, "dedupe.retention: required"}, {"dedupe.retention not a duration", FileConfig, configJSON(`{"dedupe": {"retention": "30d"}}`), `dedupe.retention: must be a duration such as "720h"`}, {"dedupe.retention a number", FileConfig, configJSON(`{"dedupe": {"retention": 3600}}`), "retention"}, {"dedupe.retention negative", FileConfig, configJSON(`{"dedupe": {"retention": "-1h"}}`), "dedupe.retention: must not be negative"}, @@ -373,6 +372,30 @@ func TestValidate_DedupeRetentionAccepted(t *testing.T) { } } +// configJSONWithout is the seed config.json less dedupe.retention. +func configJSONWithout(t *testing.T) string { + t.Helper() + var doc map[string]map[string]json.RawMessage + require.NoError(t, json.Unmarshal([]byte(configJSON(`{}`)), &doc)) + delete(doc["dedupe"], "retention") + out, err := json.Marshal(doc) + require.NoError(t, err) + return string(out) +} + +// dedupe.retention may be left out: the tenant keeps ids forever, and a +// table override may still set one. +func TestValidate_DedupeRetentionOptional(t *testing.T) { + t.Parallel() + files := validFiles() + files[FileConfig] = configJSONWithout(t) + require.NotContains(t, files[FileConfig], "retention") + doc, findings := ValidateDir(writeDir(t, files)) + require.NotNil(t, doc, "findings: %s", findingStrings(findings)) + assert.False(t, HasErrors(findings), "findings: %s", findingStrings(findings)) + assert.Nil(t, doc.Config.Dedupe.Retention) +} + // TestValidate_ClickHouseTLSPathsAreNotOpened pins that the tls block is // checked for shape only: Validate is pure and also runs on the control // plane, so paths that exist nowhere still validate, and the values reach From 76b6a9e23d8055db8e6dc9c76b34b4e0e8cc8c15 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:32:49 -0400 Subject: [PATCH 09/22] docs(settings): name dedupe.retention as the one compiled default Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- config.yaml | 4 ++-- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/api/settings_test.go | 2 +- internal/app/wire.go | 3 ++- internal/settings/settings.go | 5 +++-- internal/settings/store.go | 6 +++--- 8 files changed, 14 insertions(+), 12 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b81b90315..8806ebbb9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, deleted by the retention sweep below ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so a table name is no longer refused for the bytes it holds: one holding a NUL byte dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). +- **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate.go` (+ tests), `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, deleted by the retention sweep (see Added) ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md)). The key carries the tenant's and the table's lengths rather than separators between them, so a table name is no longer refused for the bytes it holds: one holding a NUL byte dedupes in a keyspace of its own like any other. New metrics: `wavehouse_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes, stored as its SHA-256). - **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/config.yaml b/config.yaml index 358124171..8675e2de6 100644 --- a/config.yaml +++ b/config.yaml @@ -67,8 +67,8 @@ auth: # query.default_max_rows / timestamp_bucket_seconds, # schema.refresh_interval, stream keepalive_interval / keepalive_buckets / # gap_window_minutes, mq.max_bytes_gb, cors.allowed_origins — and every key -# is required: the -# binary has no compiled defaults, so what's adopted is exactly what the +# is required except dedupe.retention (missing = "0", forever): the binary +# has no other compiled default, so what's adopted is exactly what the # files say. The server validates the directory at boot (invalid or missing # refuses to start) and reloads it on SIGHUP, on file change, or via # POST /v1/ops/settings/reload; a reload that fails validation keeps the diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e9a581e88..6ab56b3d7 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -188,7 +188,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi - **store.go** — `Store` is a passive holder: one tenant's adopted document behind an atomic pointer, swapped by the registry, which stamps it with the id of the tenant it created the store for (`Tenant()`, how a handler names its tenant to a per-tenant resource). Consumers read typed accessors per call (`ClickHouse()`, `Auth()`, `DedupeFor(table)`, `DLQFor(table)`, `Keepalive()`, …) rather than holding values. - **registry.go** — `Registry` maps a tenant id to its `Store` and owns everything that changes one. `Open` validates and adopts at boot; `Reload` re-validates the whole directory and `ReloadTenant` one tenant's folder, serialized with each other; `AfterAdopt` hooks run after every reload the registry applied, with the tenants it adopted — none when it only rejected or removed one, which a consumer holding a resource per tenant needs to hear of too; `For(id)` and `All()` see only the tenants being served, `Known()` every tenant held, rejected ones included, and `Resolve(id)` tells a rejected tenant from an unknown one. The shape is fixed at `Open`. Flat: an invalid directory refuses boot, and a rejected reload keeps the previous snapshot. Nested: fail closed per tenant — a folder with an error finding stops being served (the store keeps its document for requests already admitted, and gets the next good one) while the rest carry on; a whole-directory reload mirrors the folders, down to none (an emptied directory is not a change of shape); and a finding about the directory itself refuses boot or rejects the reload whole, leaving every tenant as it was. The tenant map is replaced whole by a reload, so a lookup is one lock-free load. - **watch.go** — `Registry.Watch`, which `internal/app` starts for a flat directory only: fsnotify on the *directory* (not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost), debounced into one reload; reloads once as soon as the watch exists so an edit between the boot read and the watch is never missed. `SIGHUP` and the reload endpoint funnel through the same serialized `Reload`. -- **seed.go** / **seed/** — The embedded (`go:embed`) starter directory with every key at its default. The binary carries no compiled defaults: `wavehouse bootstrap [dir]` writes this seed, and the compose stack and e2e fixture ship copies of it. +- **seed.go** / **seed/** — The embedded (`go:embed`) starter directory with every key at its default. The binary carries no compiled defaults except that a missing `dedupe.retention` means `"0"`: `wavehouse bootstrap [dir]` writes this seed, and the compose stack and e2e fixture ship copies of it. ### `tenant/` — Tenant Identifier diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 55a8e0d9c..9cb52b4d3 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -11,7 +11,7 @@ Boot config — the YAML file and `WH_*` environment variables on the [Configura The settings directory holds WaveHouse's file-based settings as exactly four JSON documents: [`roles.json`](#rolesjson), [`policies.json`](#policiesjson), [`pipes.json`](#pipesjson), and [`config.json`](#configjson-keys). The files are the only write path — standalone, you edit them on the host; on WaveHouse Cloud the control plane writes them — and there is no API that writes back to them. Every file must exist (an empty document is `{}` — a missing file always means deletion or a wrong path, never "defaults"), and any other entry in the directory is an error, so a typoed filename or a stray backup fails loudly instead of being silently ignored. Dot-prefixed entries are the one exception: editor swap files and the `..data` machinery Kubernetes ConfigMap mounts publish through are ignored. -Create one with `wavehouse bootstrap [dir]`: it writes all four files with every key at its default and refuses a non-empty directory, so an existing settings directory is never overwritten. The binary carries no compiled defaults — the seed is the one place they live, and what the server adopts is exactly what the files say. The seed ships no policy (`policies.json` is `{}`, `roles.json` and `pipes.json` are empty lists), so a freshly bootstrapped directory boots fail-closed — every request is denied until you write a policy. The container images ship no settings directory: they preset `WH_SETTINGS_DIR=/app/settings` and expect a bind mount there — a host directory you wrote with `bootstrap` (the reference compose file mounts the checked-in `deployments/compose/settings/`, the seed with `clickhouse.addr` pointed at the `clickhouse` service and a permissive `public` trial policy in `policies.json` / `roles.json`). A bind mount, not a named volume: the images are distroless, with no shell to edit files inside a volume. A missing mount refuses to boot rather than running on defaults nobody chose. +Create one with `wavehouse bootstrap [dir]`: it writes all four files with every key at its default and refuses a non-empty directory, so an existing settings directory is never overwritten. The binary carries no compiled defaults but one — a missing `dedupe.retention` means `"0"`, forever — so the seed is where the defaults live, and what the server adopts is exactly what the files say. The seed ships no policy (`policies.json` is `{}`, `roles.json` and `pipes.json` are empty lists), so a freshly bootstrapped directory boots fail-closed — every request is denied until you write a policy. The container images ship no settings directory: they preset `WH_SETTINGS_DIR=/app/settings` and expect a bind mount there — a host directory you wrote with `bootstrap` (the reference compose file mounts the checked-in `deployments/compose/settings/`, the seed with `clickhouse.addr` pointed at the `clickhouse` service and a permissive `public` trial policy in `policies.json` / `roles.json`). A bind mount, not a named volume: the images are distroless, with no shell to edit files inside a volume. A missing mount refuses to boot rather than running on defaults nobody chose. Check a directory with `wavehouse validate [dir]`. Both commands resolve the directory the same way — the argument, falling back to `WH_SETTINGS_DIR`, and a usage error (exit `2`) with neither — so the path you seed is the path you validate, and inside the container images (which preset `WH_SETTINGS_DIR=/app/settings`) both work with no argument at all. `validate` validates without starting the server (JSON syntax including unknown fields and duplicate keys, per-file shape rules including the required keys, and cross-file role references), prints every finding in one pass, and exits `0` for valid (warnings allowed), `1` for invalid, `2` for usage — so operators and CI can gate a settings change before it reaches a running instance. diff --git a/internal/api/settings_test.go b/internal/api/settings_test.go index 24fa80f12..462cbc8e6 100644 --- a/internal/api/settings_test.go +++ b/internal/api/settings_test.go @@ -16,7 +16,7 @@ import ( "github.com/stretchr/testify/require" ) -// fullConfig is a complete config.json (every key is required) with the +// fullConfig is a complete config.json (every key set) with the // given query.default_max_rows. func fullConfig(maxRows int) string { return fmt.Sprintf(`{"clickhouse": {"addr": "localhost:9000", "http_port": 8123, "http_scheme": "http", "database": "default", "username": "default", "query_timeout": 30, "tls": {"enabled": false, "ca_file": "", "cert_file": "", "key_file": "", "insecure_skip_verify": false, "server_name": ""}, "headers": {}, "max_open_conns": 10, "max_idle_conns": 5}, "auth": {"jwks_url": "", "role_claim": "role"}, "dedupe": {"enabled": false, "id_field": "event_id", "require_id": false, "retention": "0"}, "dlq": {"enabled": true}, "query": {"default_max_rows": %d, "timestamp_bucket_seconds": 60}, "schema": {"refresh_interval": 60}, "stream": {"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": 15}, "mq": {"max_bytes_gb": 1}, "cors": {"allowed_origins": ["*"]}}`, maxRows) diff --git a/internal/app/wire.go b/internal/app/wire.go index ac494bea7..2678646a1 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -52,7 +52,8 @@ func withoutContext(release func() error) func(context.Context) error { // configuration (dedupe, dlq, query, schema, stream, cors — see // settings.TenantConfig). Required: config.Validate already rejected an // empty settings.dir, and an invalid directory refuses boot. The binary -// carries no compiled defaults; `wavehouse bootstrap` writes the seed. A +// carries no compiled defaults but a missing dedupe.retention ("0"); +// `wavehouse bootstrap` writes the seed. A // *reload* of an invalid directory merely keeps the previous snapshot. A // nested directory (one folder per tenant, #583) fails closed per tenant // instead, at boot and on reload alike: see settings.Registry. diff --git a/internal/settings/settings.go b/internal/settings/settings.go index eeb03f358..4d5fa7a64 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -65,8 +65,9 @@ type PipesFile struct { // `cache.l1_max_cost`), listeners, the observability // exporters — and the secrets (`clickhouse.password`, `auth.jwt_secret`, // `auth.operator_key`), which never belong in a tracked JSON file. Every -// block and every top-level key inside it is REQUIRED: the binary carries no -// compiled defaults, so the adopted snapshot is exactly what the files say. +// block and every top-level key inside it is REQUIRED, but dedupe.retention +// (missing means "0", forever): the binary carries no other compiled default, +// so the adopted snapshot is exactly what the files say. // Defaults live in the seed directory (see Seed) that `wavehouse // bootstrap` writes. The fields are pointers only so Validate can tell // "absent" from the zero value and report it by path. diff --git a/internal/settings/store.go b/internal/settings/store.go index d707451d3..c2609276f 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -16,9 +16,9 @@ import ( // accessors below each resolve from a single snapshot load, so a reload lands // between lookups, never inside one. // -// There are no compiled defaults here on purpose: every key is required by -// Validate, so the snapshot is exactly what the files said when they were -// adopted. Defaults live in the seed directory (Seed / WriteSeed). +// There are no compiled defaults here on purpose, but one: every key but +// dedupe.retention (missing means "0", forever) is required by Validate, so +// the snapshot is exactly what the files said when they were adopted. Defaults live in the seed directory (Seed / WriteSeed). type Store struct { // tenant is the id the Registry created the store for; the zero value // only for a Store built outside a Registry (tests). From bea9e70150400f08f30d4b6bd29dd78f92dc3fce Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 12:34:42 -0400 Subject: [PATCH 10/22] docs(settings): the last every-key-required comment; rewrap Store's Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/settings/store.go | 3 ++- internal/settings/validate_test.go | 5 +++-- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/internal/settings/store.go b/internal/settings/store.go index c2609276f..a03546411 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -18,7 +18,8 @@ import ( // // There are no compiled defaults here on purpose, but one: every key but // dedupe.retention (missing means "0", forever) is required by Validate, so -// the snapshot is exactly what the files said when they were adopted. Defaults live in the seed directory (Seed / WriteSeed). +// the snapshot is exactly what the files said when they were adopted. +// Defaults live in the seed directory (Seed / WriteSeed). type Store struct { // tenant is the id the Registry created the store for; the zero value // only for a Store built outside a Registry (tests). diff --git a/internal/settings/validate_test.go b/internal/settings/validate_test.go index 7af9214d3..96e1ff60a 100644 --- a/internal/settings/validate_test.go +++ b/internal/settings/validate_test.go @@ -33,8 +33,9 @@ func validFiles() map[string]string { // configJSON returns the seed config.json with patch merged over it, one // level deep (a patched block's keys replace the seed's, the rest of the -// block is kept). Every key is required, so tests that care about one key -// build a complete document from the seed rather than repeating all of them. +// block is kept). Every key but dedupe.retention is required, so tests that +// care about one key build a complete document from the seed rather than +// repeating all of them. func configJSON(patch string) string { seed, err := Seed() if err != nil { From 5d77211de17e38810cbbf3335a5fd99012fedbaa Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:37:28 -0400 Subject: [PATCH 11/22] test(dedupe): a legacy value under a current key reads as absent The version-0 read test planted only a tenant-NUL-id key, which the text layout never looks up, so it passed whatever Reserve made of a legacy value. It now also plants a bare legacy id that spells a current key and checks it reads as absent and is overwritten by the commit. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/deployment.md | 2 +- internal/dedupe/embedded_test.go | 20 ++++++++++++++------ internal/dedupe/sweep.go | 2 +- 3 files changed, 16 insertions(+), 8 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index ca23f0fe6..6c103c0df 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -419,7 +419,7 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres ## Upgrading across the dedupe key change -The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys are never read, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated. The old keys never count as seen, and the dedupe sweep deletes them: its first pass runs about a minute after the instance opens, and `wavehouse_dedupe_swept_keys_total{reason="version_0"}` counts them ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)). Pebble returns their disk space as it compacts, not at once. Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. The same release adds an optional **`dedupe.retention`** key. No upgrade step is needed: a `config.json` without it keeps every id forever, as before. See [Deduplication](/settings-directory#deduplication) for a finite one. diff --git a/internal/dedupe/embedded_test.go b/internal/dedupe/embedded_test.go index 5015318f4..8dfd29733 100644 --- a/internal/dedupe/embedded_test.go +++ b/internal/dedupe/embedded_test.go @@ -140,17 +140,25 @@ func TestEmbedded_OpenFailure(t *testing.T) { assert.True(t, e.Open()) } -// Keys from before the table joined the key (#222) are never read: an id -// seen then is accepted once more after the upgrade, the documented cost of -// the new layout. -func TestEmbedded_VersionZeroKeysAreNotRead(t *testing.T) { +// Keys from before the table joined the key (#222) never count: an id seen +// then is accepted once more after the upgrade, the documented cost of the +// new layout. A tenant ‖ NUL ‖ id key is never looked up; a bare v0.1.0 id +// that spells a current key is, and its 8-byte value reads as absent. +func TestEmbedded_VersionZeroKeysDoNotCount(t *testing.T) { t.Parallel() e := NewEmbedded(t.TempDir()) m := switchedOn(t, e, "acme") require.NoError(t, e.db.Set([]byte("acme\x00e1"), make([]byte, 8), pebble.Sync)) - dup, err := mark(context.Background(), m, "e1") + stale := AppendKey(nil, KeyPrefix("acme"), Key{Table: "events", ID: "e2"}) + require.NoError(t, e.db.Set(stale, make([]byte, 8), pebble.Sync)) + for _, id := range []string{"e1", "e2"} { + dup, err := mark(context.Background(), m, id) + require.NoError(t, err) + assert.False(t, dup, id) + } + dup, err := mark(context.Background(), m, "e2") require.NoError(t, err) - assert.False(t, dup) + assert.True(t, dup, "the commit overwrote the stale value") } // A claim nobody commits, releases or reserves again leaves memory at the diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index 444ec178d..700186847 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -75,7 +75,7 @@ func (e *Embedded) startSweep(db *pebble.DB) (stop func()) { } // sweep makes one pass over the whole instance, deleting keys whose -// retention has ended and version-0 keys, which nothing reads: those from +// retention has ended and version-0 keys, which never count: those from // before ids were keyed by table (tenant ‖ 0x00 ‖ id, or the bare id before // that). They are told apart by value, since a bare id may be any bytes, a // current key's included: only commits are stored, and every commit has the From e52b336c3c7b41d3d66a608102093bda30f4c330 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:53:59 -0400 Subject: [PATCH 12/22] fix(dedupe): sweep holds the commit lock only to re-check and delete A sweep chunk held commitMu while it iterated to its 1,024th live key. Pebble skips point tombstones inside Next, so a chunk that started over a run of them (a tenant's version-0 block deleted by the first pass, or a table whose ids had all expired) held every Commit, and with it every deduped ingest response, for the whole run, on every pass until the run was compacted. The chunk now reads without the lock and collects the keys it would delete, then takes commitMu only to re-read each one and delete those still expired or version-0, without fsync as before. A Commit waits for at most 1,024 point reads and one unsynced batch, however many tombstones lie between the keys. The mid-chunk test splits in two, one per gap a Commit can land in: after the unlocked read, where the re-read keeps the key, and after the re-read, where the lock holds the Commit off until the delete is done. Each fails when its guard is removed. A new test races Commits against a chunk that starts over 300,000 tombstones: under the race detector the old chunk kept one waiting 241-278 ms, the new one at most 11 ms. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 2 +- internal/dedupe/embedded.go | 13 +-- internal/dedupe/sweep.go | 119 ++++++++++++++++++-------- internal/dedupe/sweep_test.go | 87 ++++++++++++++++--- 6 files changed, 172 insertions(+), 53 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7b7935f03..9c9feadbf 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,7 +29,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. - **Docs-site analytics for search, code copies, 404s, docs section, and live-demo connectivity** (`docs/src/components/DocsTracking.astro` (new), `docs/src/components/{PostHog,Footer,LiveDemo}.astro`): the site tracked its own CTAs but nothing a reader did on the way to one, so the questions that decide what to write next — what people search for and *don't* find, which snippets get copied, which dead links keep getting followed — had no data behind them. `docs_search` fires a second after the query settles rather than once per keystroke, carrying `query` and `result_count` read off Pagefind's own results message (the rendered list is capped at its page size, so counting the DOM would under-report); `result_count: 0` is the event worth having. `code_copied` (`page`, `language`) watches Expressive Code's copy buttons from the document rather than re-binding every code block on every navigation — the hero's install chip is not an EC block and keeps its own `hero_install_copied`. `docs_404` (`path`, `referrer`) turns broken inbound links into a list instead of a hunch. A `doc_section` property (the first path segment, `home` for `/`) puts every event in a docs area without each tracker carrying its own copy; it's stamped at capture time by a `before_send` hook in `posthog.init()` rather than `register()`, because a queued `register()` replays only after init has already captured the first hard-load `$pageview` — which would then carry the previous visit's persisted value — and `history_change` navigations update the URL before capture fires, so reading `location` in the hook is always current. `live_demo_connected` fires once per mount when the hero's SSE feed comes up rather than on its first row — named for what it measures (the demo backend answered), since a quiet minute on the repo is not a disengaged reader. The three site-wide trackers share one new `DocsTracking.astro` rendered from the footer (like `MermaidZoom` / `ScrollHints`) and delegate from `document`, since Pagefind, Expressive Code, and the 404 route all own their own markup — some of it created after page load. -- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains an optional **`dedupe.retention`** key, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. A `config.json` without the key keeps ids forever, and a table override without one inherits the tenant's, so an existing directory needs no change. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk and deletes the expired and version-0 ones, without fsync, under a lock `Commit` also takes, so an id committed again after the sweep read it is never deleted. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. +- **Dedupe retention per tenant and table, and a sweep that deletes expired ids** (`internal/settings/{settings,validate,store}.go` (+ tests), `internal/settings/seed/config.json`, `internal/dedupe/{embedded,sweep}.go` (+ tests), `internal/api/ingest.go` (+ tests), `internal/testutil/mocks.go`, `deployments/compose/settings/config.json`, `tests/e2e/fixtures/settings/config.json`, `config.yaml`, `docs/src/content/docs/{settings-directory.mdx,deployment,durability,architecture}.md`, `AGENTS.md`): [#220](https://github.com/Wave-RF/WaveHouse/issues/220). `config.json` gains an optional **`dedupe.retention`** key, overridable per table in `dedupe.tables.
.retention`: how long a committed id stays a duplicate, as a duration string (`"720h"`), or `"0"` to keep it forever, which is the seed value and the behaviour before this release. A `config.json` without the key keeps ids forever, and a table override without one inherits the tenant's, so an existing directory needs no change. A finite retention below `"2m"`, the ingest queue's duplicate window, is refused rather than raised to the minimum: every deduped record is published under an idempotency key derived from its id, so an id re-sent after a shorter retention would be claimed again and dropped by the queue as a copy while the client was told it was accepted. It is hot-reloadable and read per record like `id_field`; a change applies to ids committed after it, and a reload landing mid-window commits each record with the retention it was prepared under. The embedded Pebble store already treated an expired id as new; it now also deletes expired keys, and the version-0 keys the key-layout change left behind, in a background sweep that starts a minute after the instance opens and repeats hourly. It reads 1,024 keys per chunk without a lock, then re-reads the expired and version-0 ones under a lock `Commit` also takes and deletes, without fsync, those that still are, so an id committed again after the sweep read it is never deleted, and a `Commit` waits for at most one chunk's re-reads, never for the deleted keys a chunk steps over. New metric `wavehouse_dedupe_swept_keys_total{reason="expired"|"version_0"}`. `settings.Store.DedupeFor` now returns a `settings.Dedupe` struct rather than three values. ### Changed diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 642d25a03..b7fa3db58 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -125,7 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose records were definitely not published (a refused or never-sent publish; one whose outcome is unknown is left to lapse instead). A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. -- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync, each value carrying its expiry (`0` = never), which `Reserve` honors on read. A background sweep (`sweep.go`), started when the instance opens and stopped before it closes, deletes expired keys and the version-0 keys from before the table joined the key — told apart by their value, which is never a current commit's, since a bare id from before tenants led the key could spell a current one ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)): a minute after opening, then hourly, 1,024 keys per chunk, holding a lock `Commit` also takes, so a key re-committed after the sweep read it is never deleted; `wavehouse_dedupe_swept_keys_total{reason}` counts what it deletes. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync, each value carrying its expiry (`0` = never), which `Reserve` honors on read. A background sweep (`sweep.go`), started when the instance opens and stopped before it closes, deletes expired keys and the version-0 keys from before the table joined the key — told apart by their value, which is never a current commit's, since a bare id from before tenants led the key could spell a current one ([#220](https://github.com/Wave-RF/WaveHouse/issues/220)): a minute after opening, then hourly, 1,024 keys per chunk, read without a lock, so the deleted keys a chunk steps over (Pebble keeps them until it compacts) never hold up a `Commit`, then re-read under a lock `Commit` also takes and deleted only if still expired or version-0, so a key re-committed after the sweep read it is never deleted; `wavehouse_dedupe_swept_keys_total{reason}` counts what it deletes. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 32e9f5b2e..48fc88173 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path; it reads 1,024 keys at a time, deleting the expired ones, and a commit that arrives mid-chunk waits for that chunk. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. +With a finite `dedupe.retention`, expired ids are deleted by a background sweep, an hour apart. Its deletes are not fsynced (a delete lost to a crash is redone by the next pass), so it adds no sync to the ingest path. It reads 1,024 keys at a time without holding up commits, then re-reads the expired ones and deletes those still expired; a commit waits only for that last step, at most 1,024 point reads and one unsynced write, however many deleted keys the read stepped over. An expired id is already treated as new by the next claim of it, sweep or no sweep, so retention never depends on the sweep having run. A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. For the same reason a finite `dedupe.retention` must be at least those two minutes: an id re-sent after a shorter retention ended would be claimed again, then dropped by the stream as a copy while the client was told it was accepted. Settings validation refuses one below it. diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index d0fe87ac9..5eece3f3b 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -35,8 +35,9 @@ type Embedded struct { open int // tenant stores open over db stopSweep func() // stops db's sweep - // commitMu is read-held by Commit and held by a sweep chunk, so a sweep - // never deletes a key a Commit rewrote after the sweep read it. + // commitMu is read-held by Commit and held by a sweep chunk while it + // re-reads and deletes, so a sweep never deletes a key a Commit rewrote + // after the sweep read it. commitMu sync.RWMutex sweepFirst time.Duration sweepEvery time.Duration @@ -47,9 +48,11 @@ type Embedded struct { // readHook, when set, runs before each Pebble read in Reserve; a test // makes it fail to exercise Reserve's all-or-nothing error path. readHook func() error - // sweepHook, when set, runs in a sweep chunk between reading its keys and - // deleting them; a test races a Commit into that gap. - sweepHook func() + // sweepScanHook and sweepDeleteHook, when set, run in a sweep chunk: + // between its unlocked read and its re-read, and between its re-read and + // its delete. A test races a Commit into each gap. + sweepScanHook func() + sweepDeleteHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index 700186847..f1c4ca180 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -3,6 +3,7 @@ package dedupe import ( "bytes" "context" + "errors" "fmt" "log/slog" "time" @@ -20,9 +21,9 @@ import ( const ( sweepInterval = time.Hour sweepFirstDelay = time.Minute - // sweepChunk keys are read and deleted per lock hold, with sweepPause - // between chunks: at most ~100k keys a second, and a Commit never waits - // longer than one chunk. + // sweepChunk keys are read per chunk, with sweepPause between chunks: at + // most ~100k keys a second. A Commit waits only for a chunk's re-reads + // and deletes, never for its read, however many tombstones it skips. sweepChunk = 1024 sweepPause = 10 * time.Millisecond ) @@ -98,21 +99,41 @@ func (e *Embedded) sweep(ctx context.Context, db *pebble.DB) (sweepResult, error } // sweepChunk deletes the sweepable keys among the next sweepChunk keys from -// from, returning where the next chunk starts (nil at the end). It holds -// commitMu, so no Commit lands between reading a key and deleting it: a key -// re-committed after it expired is never deleted with its new value. +// from, returning where the next chunk starts (nil at the end). It reads them +// without commitMu, since Pebble skips the tombstones between keys inside the +// read and a run of them left by an earlier pass would otherwise hold every +// Commit for its whole length. func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, res *sweepResult) ([]byte, error) { - e.commitMu.Lock() - defer e.commitMu.Unlock() - now := e.now() + candidates, next, err := sweepCandidates(db, from, e.now()) + if err != nil || len(candidates) == 0 { + return next, err + } + if e.sweepScanHook != nil { + e.sweepScanHook() + } + expired, v0, err := e.deleteSweepable(db, candidates) + if err != nil { + return nil, err + } + res.Expired += int(expired) + res.Version0 += int(v0) + if expired > 0 { + sweptKeysCounter.Add(ctx, expired, metric.WithAttributes(attribute.String(sweptAttribute, sweptExpired))) + } + if v0 > 0 { + sweptKeysCounter.Add(ctx, v0, metric.WithAttributes(attribute.String(sweptAttribute, sweptVersion0))) + } + return next, nil +} + +// sweepCandidates reads the next sweepChunk keys, starting at from, and +// returns those sweepable at now and where the next chunk starts (nil at the +// end). +func sweepCandidates(db *pebble.DB, from []byte, now time.Time) (candidates [][]byte, next []byte, err error) { it, err := db.NewIter(&pebble.IterOptions{LowerBound: from}) if err != nil { - return nil, fmt.Errorf("dedupe sweep: %w", err) + return nil, nil, fmt.Errorf("dedupe sweep: %w", err) } - b := db.NewBatch() - defer func() { _ = b.Close() }() - var next []byte - var expired, v0 int64 seen := 0 for valid := it.First(); valid; valid = it.Next() { if seen == sweepChunk { @@ -120,38 +141,66 @@ func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, r break } seen++ - k := it.Key() - val := it.Value() - switch { - case !isCommit(val): + if sweepReason(it.Value(), now) != "" { + candidates = append(candidates, bytes.Clone(it.Key())) + } + } + if err := it.Close(); err != nil { + return nil, nil, fmt.Errorf("dedupe sweep: %w", err) + } + return candidates, next, nil +} + +// deleteSweepable re-reads each candidate and deletes those still sweepable, +// holding commitMu so no Commit lands between the re-read and the delete: a +// key re-committed after the unlocked read is never deleted with its new +// value. +func (e *Embedded) deleteSweepable(db *pebble.DB, candidates [][]byte) (expired, v0 int64, err error) { + e.commitMu.Lock() + defer e.commitMu.Unlock() + now := e.now() + b := db.NewBatch() + defer func() { _ = b.Close() }() + for _, k := range candidates { + val, closer, err := db.Get(k) + if errors.Is(err, pebble.ErrNotFound) { + continue + } + if err != nil { + return 0, 0, fmt.Errorf("dedupe sweep: %w", err) + } + reason := sweepReason(val, now) + _ = closer.Close() + switch reason { + case sweptVersion0: v0++ - case committedExpired(val, now): + case sweptExpired: expired++ default: continue } if err := b.Delete(k, nil); err != nil { - _ = it.Close() - return nil, fmt.Errorf("dedupe sweep: %w", err) + return 0, 0, fmt.Errorf("dedupe sweep: %w", err) } } - if err := it.Close(); err != nil { - return nil, fmt.Errorf("dedupe sweep: %w", err) - } - if e.sweepHook != nil { - e.sweepHook() + if e.sweepDeleteHook != nil { + e.sweepDeleteHook() } // NoSync: a delete lost to a crash is redone by the next pass. if err := b.Commit(pebble.NoSync); err != nil { - return nil, fmt.Errorf("dedupe sweep: %w", err) - } - res.Expired += int(expired) - res.Version0 += int(v0) - if expired > 0 { - sweptKeysCounter.Add(ctx, expired, metric.WithAttributes(attribute.String(sweptAttribute, sweptExpired))) + return 0, 0, fmt.Errorf("dedupe sweep: %w", err) } - if v0 > 0 { - sweptKeysCounter.Add(ctx, v0, metric.WithAttributes(attribute.String(sweptAttribute, sweptVersion0))) + return expired, v0, nil +} + +// sweepReason is why the sweep deletes a key holding val at now, or "" when +// it keeps it. +func sweepReason(val []byte, now time.Time) string { + switch { + case !isCommit(val): + return sweptVersion0 + case committedExpired(val, now): + return sweptExpired } - return next, nil + return "" } diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index da4ec9512..0c7ff46e9 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -100,28 +100,52 @@ func TestEmbedded_SweepDeletesExpiredAndVersionZeroKeys(t *testing.T) { assert.Equal(t, sweepResult{}, res, "a second pass finds nothing") } -// A Commit that arrives while a sweep chunk has read an expired key but not -// yet deleted it waits for the chunk, so the new commit is never deleted with -// the old value. Without the lock the Commit lands in the gap and the sweep -// then deletes it; the wait below only ever lets that pass, never fail. -func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { - t.Parallel() +// expiredAndClaimed commits id "e1" with an hour's retention, lets it expire +// and claims it again, for a test to commit mid-sweep. +func expiredAndClaimed(t *testing.T) (*Embedded, *Managed, []Claim) { + t.Helper() e := NewEmbedded(t.TempDir()) clock := newStepClock() SetClock(e, clock.now) m := switchedOn(t, e, "acme") commitIDs(t, m, time.Hour, "e1") clock.advance(2 * time.Hour) - claims, err := m.Reserve(context.Background(), []Key{{Table: "events", ID: "e1"}}, DefaultLease) require.NoError(t, err) require.Equal(t, Claimed, claims[0].Status, "expired: claimable again") + return e, m, claims +} + +// A key committed again after a sweep chunk read it as expired, but before +// the chunk re-read it, is kept: the re-read sees the new commit. +func TestEmbedded_SweepKeepsAKeyCommittedAfterItsRead(t *testing.T) { + t.Parallel() + e, m, claims := expiredAndClaimed(t) + var commitErr error + e.sweepScanHook = func() { commitErr = m.Commit(context.Background(), claims, time.Hour) } + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + require.NoError(t, commitErr) + assert.Equal(t, sweepResult{}, res) + + dup, err := mark(context.Background(), m, "e1") + require.NoError(t, err) + assert.True(t, dup, "the commit made after the read survived the sweep") +} + +// A Commit that arrives while a sweep chunk has re-read an expired key but +// not yet deleted it waits for the chunk, so the new commit is never deleted +// with the old value. Without the lock the Commit lands in the gap and the +// sweep then deletes it; the wait below only ever lets that pass, never fail. +func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { + t.Parallel() + e, m, claims := expiredAndClaimed(t) done := make(chan error, 1) - e.sweepHook = func() { + e.sweepDeleteHook = func() { go func() { done <- m.Commit(context.Background(), claims, time.Hour) }() select { - case <-done: - done <- nil + case err := <-done: + done <- err case <-time.After(50 * time.Millisecond): } } @@ -135,6 +159,49 @@ func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { assert.True(t, dup, "the commit made mid-chunk survived the sweep") } +// A chunk that starts over a long run of tombstones, as a tenant's version-0 +// block leaves until Pebble compacts it, reads through the run without the +// lock, so a Commit racing it waits for the chunk's re-reads and deletes +// alone. Sized for the race detector, which the unit suite runs under: there, +// a chunk holding the lock across this run kept a Commit waiting ~250 ms. +func TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits(t *testing.T) { + t.Parallel() + e := NewEmbedded(t.TempDir()) + m := switchedOn(t, e, "acme") + b := e.db.NewBatch() + for i := range 300_000 { + require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%06d", i), nil)) + } + // A version-0 key after the run, so the chunk has one to delete. + require.NoError(t, b.Set([]byte("acme\x01"), make([]byte, 8), nil)) + require.NoError(t, b.Commit(pebble.NoSync)) + require.NoError(t, e.db.Flush()) + + type swept struct { + res sweepResult + err error + } + done := make(chan swept, 1) + go func() { + res, err := e.sweep(context.Background(), e.db) + done <- swept{res, err} + }() + var slowest time.Duration + for i := 0; ; i++ { + select { + case s := <-done: + require.NoError(t, s.err) + require.Equal(t, sweepResult{Version0: 1}, s.res) + assert.Less(t, slowest, 100*time.Millisecond, "the slowest Commit racing the sweep") + return + default: + } + start := time.Now() + commitIDs(t, m, 0, fmt.Sprintf("c%d", i)) + slowest = max(slowest, time.Since(start)) + } +} + // A retention is honoured on read before any sweep has run: the key is a // duplicate until the retention ends and claimable from that instant. func TestEmbedded_RetentionHonouredOnRead(t *testing.T) { From ffa4dd6875c168154e68fd4219601d7b5a380ea3 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:54:27 -0400 Subject: [PATCH 13/22] docs(settings): name dedupe.retention in the last no-defaults claims dedupe.retention is the one config.json key the binary defaults (missing means "0", forever). The ingest handler's dedupe comment, the seed's doc comment, a registry test comment and the settings-directory CHANGELOG entry still said there were none. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- internal/api/ingest.go | 4 ++-- internal/settings/registry_test.go | 2 +- internal/settings/seed.go | 3 ++- 4 files changed, 6 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9c9feadbf..8019c2e34 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -24,7 +24,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Schema discovery captures each table's DDL, its columns' ordinals and default expressions, and the server version** (`internal/discovery/discovery.go`, `internal/testutil/testutil.go`): `Column` gains `DefaultExpression` and `Position` (both from a widened `system.columns` select), `TableSchema` gains `DDL` from `system.tables.create_table_query`, and `SchemaRegistry` gains `ServerVersion()` from a `SELECT version()` probe next to the existing `SELECT timezone()`. Groundwork for the native type layer, captured on the same refresh as the columns so a stale version cannot outlive the schemas it describes. That is a publication guarantee, not a same-server one: `chconn.Manager` resolves the connection per call, so a reload changing `clickhouse.addr` mid-refresh can still pair a version from one server with schemas from another — narrow, and self-correcting on the next refresh. `DDL` is `json:"-"` and does **not** appear in `/v1/ops/schema`: that endpoint marshals `TableSchema` straight to the client, and an external-engine table (S3, MySQL, PostgreSQL, Kafka) renders its wiring there unconditionally — endpoint, bucket or host, database, username, S3 access key id. ClickHouse masks the password itself as `[HIDDEN]` from ~23.9 (verified on 26.7.3), so the exposure is the topology rather than the secret — except on an older server, or one with `display_secrets_in_show_and_select` enabled. `position` and `default_expression` are additive fields in the response. A table listed in `system.tables` with no `system.columns` rows is skipped rather than published column-less, and both new queries fail the refresh on error exactly as `timezone()` and `system.columns` do — callers keep the prior cache and retry. -- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the sweeper purges it back under the limit, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. +- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the sweeper purges it back under the limit, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** but one — every `config.json` key except `dedupe.retention` (missing means `"0"`, forever) is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. - **"Was this page helpful?" feedback widget on every docs page** (`docs/src/components/PageFeedback.astro` (new), `docs/src/components/Footer.astro`): a thumbs-up / thumbs-down vote below the page content, captured to PostHog as `docs_feedback` with `{ helpful, page }`. It renders from `Footer.astro`'s sidebar branch — the same indirection the Cloud CTA uses — rather than a per-page import or frontmatter flag, so every content page gets it automatically, including ones not written yet; it sits *below* the Cloud CTA on the pages that carry one, and splash pages (the homepage and 404) take the other footer branch and never render it. One vote per page per visitor: the choice is remembered in `localStorage` keyed by pathname, and a revisit renders the thanks message instead of re-prompting (storage is a nicety, not the record — a browser with storage disabled still votes). - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 1189547a4..58126f942 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -704,8 +704,8 @@ func (h *IngestHandler) prepareRecord( // Optional deduplication. The dedupe settings resolve per record // from one snapshot (table override → global; the settings directory - // always states them, so no compiled fallback is needed), so a reload - // lands at a record boundary. A Deduplicator without a settings source is + // states them all but dedupe.retention, whose absence means "0"), so a + // reload lands at a record boundary. A Deduplicator without a settings source is // a wiring bug, not a mode — main wires both or neither. The id is claimed // in ingestWindow, once every record of the window is encoded, so nothing // but the publish can fail while the claim is held. diff --git a/internal/settings/registry_test.go b/internal/settings/registry_test.go index 3ac222ccb..a5abce849 100644 --- a/internal/settings/registry_test.go +++ b/internal/settings/registry_test.go @@ -79,7 +79,7 @@ func TestRegistry_ReloadWithWarningsAdopts(t *testing.T) { // TestOpen_RejectsInvalid pins the boot contract: an invalid directory yields // no Registry at all — there is no "store without a document" state and no -// compiled defaults to fall back on. +// compiled default for a required key to fall back on. func TestOpen_RejectsInvalid(t *testing.T) { t.Parallel() files := validFiles() diff --git a/internal/settings/seed.go b/internal/settings/seed.go index d9e979ae5..5315c1730 100644 --- a/internal/settings/seed.go +++ b/internal/settings/seed.go @@ -10,7 +10,8 @@ import ( // seedFS holds the starter settings directory: every file present, every // key set to its default. The checked-in seed/ directory is the ONE place -// defaults live — the binary has no compiled fallbacks. Its one consumer is +// defaults live — the binary's one compiled fallback is dedupe.retention +// (missing means "0", forever). Its one consumer is // this embed, so `wavehouse bootstrap` can write the directory anywhere // without a source tree; the container images ship no settings (the operator // mounts or seeds /app/settings), same as they ship no policy file. From 860d92985818df094541427f876a69ecbd48477b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:20:22 -0400 Subject: [PATCH 14/22] test(mq): state the lease/duplicate-window invariant as 2*lease+1s The in-flight 503 for an uncertain publish sends the FULL dedupe lease as Retry-After, so an obedient client's retry can land up to ~2*lease after the original Reserve -- not just one lease later -- and a claim's expiry can itself round up by up to a second on some backends (DynamoDB, for one). "lease <= window" understates what the embedded queue's duplicate window actually has to cover. Pin the real invariant in TestIngest_DedupeLeaseFitsTheDuplicateWindow (2*dedupe.DefaultLease + time.Second <= mq.EmbeddedDuplicateWindow), correct the embedded.go comment that said "must not exceed", and reword the same claim in durability.md, api.md, architecture.md and CHANGELOG.md. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 ++-- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/durability.md | 2 +- internal/api/ingest_window_test.go | 12 ++++++++---- internal/mq/embedded.go | 28 ++++++++++++++++++++------- 6 files changed, 34 insertions(+), 16 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index dac04cfc0..2fddb0e3f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -86,7 +86,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **Dedupe claims an id, publishes, then commits it — and keys it by tenant, table and id** (`internal/dedupe/{dedupe,key,embedded,managed}.go` (+ tests), `internal/dedupe/dedupetest/` (new), `internal/api/ingest.go` (+ tests), `internal/settings/validate_test.go`, `internal/keyenc/keyenc.go`, `internal/testutil/mocks.go`, `internal/app/app_test.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment,development}.md`, `settings-directory.mdx`, `sdk/reference.md`): `CheckAndMark` is replaced by a two-phase `Reserve` → `Commit` / `Release` contract with a lease on the pending claim, and every backend now runs one conformance suite. Four bugs go with it. Two concurrent requests carrying one id no longer both publish it: Pebble's check and claim happen under one lock, and the loser answers `503` with `Retry-After` while the winner is still publishing ([#390](https://github.com/Wave-RF/WaveHouse/issues/390)). A publish that fails gives its id back, so the retry a `503` asks for is published instead of skipped as a duplicate of a record that never reached the queue ([#384](https://github.com/Wave-RF/WaveHouse/issues/384)). The same id in two tables is two ids ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s keyspace half). An explicit `null` id is a missing id — rejected under `require_id`, published un-deduped otherwise — instead of the one id `""` that made every null record after the first a duplicate ([#370](https://github.com/Wave-RF/WaveHouse/issues/370)). **Upgrade:** the key layout changes, so an id seen before the upgrade is accepted once more after it; nothing is migrated, and the old keys are left in `/pebble`, unread, with nothing removing them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep) ([Deployment → Upgrading across the dedupe key change](https://github.com/Wave-RF/WaveHouse/blob/main/docs/src/content/docs/deployment.md#upgrading-across-the-dedupe-key-change)). The key is readable text, `/
/` (for example `acme/clicks/evt-123`), with the table and id escaped and joined by `internal/keyenc`, the escaping NATS subject tokens already use, so any table name gets a keyspace of its own, including one holding a NUL byte or a `/`. New metrics: `wavehouse_ingest_dedupe_commit_failed_total` (a published record whose id failed to commit; the claim lapses with its lease) and `wavehouse_dedupe_hashed_id_total` (an id over 1,024 bytes once escaped, stored as its SHA-256). -- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, which the 30-second lease fits inside. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_ingest_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. +- **Ingest runs in windows of 256 records over the dedupe contract, and a dedupe store that cannot answer is a `503`** (`internal/api/ingest.go` (+ tests), `internal/mq/{mq,embedded}.go` (+ tests), `internal/dedupe/key.go` (+ tests), `internal/testutil/mocks.go`, `AGENTS.md`, `docs/src/content/docs/{api,architecture,durability}.md`, `settings-directory.mdx`, `sdk/reference.md`): each window of a request is prepared, then reserved in one dedupe call, published in order, and committed in one call, so a batch costs one dedupe round trip per phase per window rather than per record — on Pebble, one commit `fsync` per window (a 1,000-record batch: four syncs instead of a thousand, 24 ms against 5.7 s of dedupe time measured with the queue stubbed). Every deduped record is published under an idempotency key (`mq.WithIdempotencyKey`, JetStream's message id, derived by `dedupe.IdempotencyKey`), and each tenant's ingest stream now keeps an explicit two-minute duplicate window, sized to `2 × the 30-second lease + 1s`: an uncertain publish's `503` sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and the `+1s` covers a backend whose claim expiry itself rounds up by that much. That closes the last path of [#384](https://github.com/Wave-RF/WaveHouse/issues/384): a publish that fails with an unknown outcome (anything but a full queue) keeps its record's claim until the lease lapses instead of releasing it, and a retry after the lease but within two minutes of the first publish is dropped by the queue if the first copy was stored (a later one is stored again). A dedupe store that is not open or that reports `dedupe.ErrUnavailable` now answers `503 {"error":"dedupe store unavailable"}` with `Retry-After: 5`, which the SDK retries, rather than `500 dedupe failed`. A mid-body read error or dedupe failure now drops the open window unpublished, where records before it used to be published; `wavehouse_ingest_dedupe_commit_failed_total` and `wavehouse_ingest_dedupe_disabled_total` count records, as before, now added a window at a time. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index efdd4373e..fd8d37c01 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -295,10 +295,10 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 500 | `{"error":"dedupe failed"}` | Deduplication backend error | | 503 | `{"error":"dedupe store unavailable"}` | Dedupe is on and its store is not open (for example, it failed to open on a reload); `Retry-After: 5`. Nothing was published, so the retry is safe | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | -| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown: the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue remembers the key for two minutes after the first publish, so a retry inside that window stores no second copy (the SDK's, after the 30-second `Retry-After`, lands inside it); a later one is stored again. | +| 500 | `{"error":"publish failed"}` | Message queue error whose outcome is unknown, other than a full queue or an unreachable broker (below): the event may have been stored. With dedupe on, the record's id is left to lapse with the dedupe lease (30 seconds) rather than given back: a retry inside the lease answers the in-flight `503`, and one after it is published under the same idempotency key, which the queue drops if the first copy was stored. The queue's duplicate window (two minutes) covers up to ~2×lease plus a margin, not just the lease itself, so a retry timed off `Retry-After` anywhere in this flow stores no second copy; a much later one is stored again. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | | 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | -| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time (a transient broker failure, not a full queue). Response includes `Retry-After: 5` header. Reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. With dedupe on, the record's id is given back so the retry can publish — but an unavailable broker that timed out may already have stored the event, so that retry can publish a second copy (the windowed-ingest follow-up closes this with an idempotency key). | +| 503 | `{"error":"service unavailable"}` | The message queue could not be reached or did not answer in time (`mq.ErrUnavailable`, a transient broker failure, not a full queue) — reserved for an external broker ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)): the embedded broker never reports this, and its publish failures are the `500` above. As for the `500`, the record's id is left to lapse rather than given back, so a retry cannot land as a second copy; `Retry-After` is that lease, rounded up to whole seconds, when dedupe was on for the record, else the flat `Retry-After: 5`. | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 15ab9574f..906469bdd 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -164,7 +164,7 @@ The **only** package that imports NATS/JetStream — a `depguard` rule in `.gola - **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: ASCII letters, digits, `_` and `-` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **deadletter.go** — `deadLetterTables`, the per-table count `DeadLetterCounts` reports: a dead-letter stream's per-subject counts, each subject parsed back to its topic and counted under its table — every scope of a table under the table itself, so a dotted table name never shares a count with a table + scope pair — and a table filter keeps that table with all of its scopes. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`) and remembering idempotency keys for `EmbeddedDuplicateWindow` (two minutes, which a dedupe lease must not exceed), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`) and remembering idempotency keys for `EmbeddedDuplicateWindow` (two minutes, sized to `2 × the dedupe lease + 1s` — the in-flight `503` sends the full lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original `Reserve`, and the `+1s` covers a backend whose claim expiry itself rounds up by that much), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. - **mqtest/** — The conformance suite for `Broker` (`mqtest.Run`): the behavior the rest of the process relies on — publish and consume round trips with names that need encoding, per-tenant order, redelivery, dead-lettering and its counts, replay bounds and isolation, the one `failed` report of a consumer whose delivery ends underneath it — checked through the interfaces alone, with no stream or subject name in sight. Each implementation runs it from a test of its own — the embedded one from `mqtest/embedded_test.go`, a test binary apart from `internal/mq`'s so the two share no 15s budget — handing it a fresh broker per case and flags (`mqtest.Caps`) for the few places where backends legitimately differ: whether a full queue refuses its own tenant alone, whether `PurgeAcked` removes anything, whether a tenant never given a budget has a dead-letter queue to report on, and whether `CreateConsumer` configures the durable or only finds one. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index e6a28c58a..324e1db10 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -62,7 +62,7 @@ The tell for a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host) With [deduplication](/settings-directory#deduplication) on, a `200` also means the records' ids were committed to the dedupe store, or, if that commit failed, that the failure was counted by `wavehouse_ingest_dedupe_commit_failed_total` and the ids lapse with their lease. On the embedded Pebble store that commit is an `fsync` of its own. It is taken once per window of up to 256 records of a request, after the window's publishes, rather than once per record: a 1,000-record batch costs four dedupe syncs, not a thousand. Measured with `BenchmarkIngest_DedupBatchOnPebble` on a developer laptop, with the queue stubbed out so only the dedupe store touched disk, the dedupe work for that batch took 24 ms windowed against 5.7 s one record at a time; the JetStream publishes' own fsyncs come on top. A single-record request still pays one sync for its publish and one for its commit. -A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The 30-second lease sits well inside those two minutes, so a prompt retry is covered. +A publish can also fail after JetStream stored the event (a timeout on the ack). The record's id is then left to lapse with its 30-second dedupe lease rather than given back, and every deduped record is published under an idempotency key derived from its tenant, table and id, which each tenant's ingest stream remembers for two minutes after the first publish. A retry after the lease but inside those two minutes is therefore dropped by the stream rather than stored twice; one later than that is stored again. The duplicate window has to cover more than the lease alone: the `503` for an uncertain publish sends the *full* lease as `Retry-After`, so an obedient client's retry can land up to ~2×lease after the original request, and a claim's expiry can itself round up by a further second on some backends — the invariant the queue configuration and its tests pin is `2×lease + 1s ≤ window`, not just `lease ≤ window`. Two minutes against a 30-second lease clears that with room to spare. ## Check your storage before you trust it diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index 3d1dc931a..ba350eedd 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -249,12 +249,16 @@ func TestIngest_Windows_OutcomesStayInOrder(t *testing.T) { } } -// The embedded queue must remember an idempotency key for at least a lease: -// the retry of an uncertain publish lands after the lease, and only the queue's -// duplicate window drops its second copy. +// The embedded queue must remember an idempotency key for at least two +// leases plus a second: the in-flight 503 of an uncertain publish sends the +// full lease as Retry-After, so a client that obeys it can republish up to +// ~2*lease after the original Reserve, and a claim's expiry can itself round +// up by up to a second (a DynamoDB backend, for one). Only the queue's +// duplicate window running at least that long guarantees it still drops the +// retry's second copy. func TestIngest_DedupeLeaseFitsTheDuplicateWindow(t *testing.T) { t.Parallel() - assert.LessOrEqual(t, dedupe.DefaultLease, mq.EmbeddedDuplicateWindow) + assert.LessOrEqual(t, 2*dedupe.DefaultLease+time.Second, mq.EmbeddedDuplicateWindow) } // faultyPublisher publishes through a real broker and fails the calls fail diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 304e1f13a..1e3853edf 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -227,6 +227,13 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { held uint64 } dlqs := map[tenant.ID]dlqState{} + // duplicates is the ingest stream's own Duplicates window as found on + // disk, keyed alongside dlqs: a stream from before EmbeddedDuplicateWindow + // existed, or reopened under a different value, must not be counted as + // already at budget below, or SetMaxBytes(same budget) short-circuits and + // the stale window is never brought forward (measured: a stream with + // Duplicates=10s kept 10s after NewEmbedded + SetMaxBytes(same budget)). + duplicates := map[tenant.ID]time.Duration{} streams := e.js.ListStreams(ctx) for info := range streams.Info() { name := info.Config.Name @@ -234,6 +241,7 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { q := e.queue(id) q.ingest = true q.asked, q.ingestCap = info.Config.MaxBytes, info.Config.MaxBytes + duplicates[id] = info.Config.Duplicates } else if id, ok := streamTenant(dlqStreamPrefix, name); ok { e.queue(id).dlq = true dlqs[id] = dlqState{limit: info.Config.MaxBytes, held: info.State.Bytes} @@ -244,14 +252,16 @@ func (e *EmbeddedNATS) takeStock(ctx context.Context) error { } // A pair is at its budget when its dead-letter stream is at a tenth of // the ingest cap, or above it holding more than that: the shrink guard's - // doing. Anything else is a pair a stop or a failed update left split, or - // one missing its dead-letter stream, so its budget stays unapplied and - // the boot's SetMaxBytes applies it to both streams again. + // doing, AND its ingest stream's duplicate window already matches + // EmbeddedDuplicateWindow. Anything else is a pair a stop or a failed + // update left split, one missing its dead-letter stream, or one whose + // duplicate window is stale, so its budget stays unapplied and the boot's + // SetMaxBytes applies it — and the current window — to both streams again. for id, q := range e.queues { d, ok := dlqs[id] tenth := q.asked / dlqShare guarded := d.limit > tenth && d.held <= math.MaxInt64 && int64(d.held) > tenth - if q.ingest && ok && (d.limit == tenth || guarded) { + if q.ingest && ok && (d.limit == tenth || guarded) && duplicates[id] == EmbeddedDuplicateWindow { q.maxBytes = q.asked } e.record(id, q) @@ -318,9 +328,13 @@ func (e *EmbeddedNATS) record(id tenant.ID, q *tenantQueue) { } // EmbeddedDuplicateWindow is how long an ingest queue remembers a -// WithIdempotencyKey key. A dedupe lease must not exceed it: a claim left to -// lapse after an uncertain publish is republished once the lease ends, and -// only this window drops that second copy. +// WithIdempotencyKey key. It must be at least 2*lease + 1s: a claim left to +// lapse after an uncertain publish is republished once the lease ends, but +// the in-flight 503 tells a client to retry only after the FULL lease, so an +// obedient client's retry can land up to ~2*lease after the original +// Reserve; the +1s covers a backend (DynamoDB, for one) that rounds a +// claim's expiry up by as much. Only a window at least that long guarantees +// this queue still drops the retry's second copy. const EmbeddedDuplicateWindow = 2 * time.Minute // ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard From 65524789d0c9d3f836b56df5859a130219d3da49 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:19 -0400 Subject: [PATCH 15/22] test(mq): pin the takeStock boot fix for a stale duplicate window The takeStock fix landed with the previous commit (embedded.go), since both are about the same staleness question; this adds its coverage. TestNewEmbedded_TakeStockRefreshesAStaleDuplicateWindow reopens a store whose ingest stream was left with a Duplicates window other than EmbeddedDuplicateWindow, then calls SetMaxBytes with the SAME budget as before and checks the window is brought forward -- the path TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat did not cover (its stream is created directly, never recorded by takeStock, so SetMaxBytes's same-budget early return never applies to it). Corrected that test's comment to say so, pointing at the new one. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/mq/embedded_test.go | 36 ++++++++++++++++++++++++++++++++++-- 1 file changed, 34 insertions(+), 2 deletions(-) diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index cf96ee9d7..7c7a87a0d 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -163,8 +163,12 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { } // A repeated idempotency key inside the duplicate window is dropped as a -// success, so an uncertain publish can be republished safely; a queue made -// with another window gets this one on its next budget apply. +// success, so an uncertain publish can be republished safely. This stream is +// created directly, never recorded by takeStock, so SetMaxBytes's next +// budget apply always runs and picks up the current window; +// TestNewEmbedded_TakeStockRefreshesAStaleDuplicateWindow covers the boot +// path, where takeStock itself must not mistake a stale window for one +// already at budget. func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { e := openEmbedded(t, t.TempDir()) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) @@ -1500,6 +1504,34 @@ func TestNewEmbedded_TakesStockOfTheQueuesOnDisk(t *testing.T) { assert.Equal(t, int64(8<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) } +// takeStock must not count a stream as at its budget when its Duplicates +// window is stale (from before EmbeddedDuplicateWindow existed, or changed +// underneath it): otherwise SetMaxBytes's same-budget early return never lets +// a later apply bring the window forward, and the stream keeps whatever it +// had indefinitely. +func TestNewEmbedded_TakeStockRefreshesAStaleDuplicateWindow(t *testing.T) { + t.Parallel() + dir := storeDir(t) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + first, err := NewEmbedded(dir) + require.NoError(t, err) + require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) + stale := ingestStreamConfig("acme", 8<<20) + stale.Duplicates = 10 * time.Second + _, err = first.js.UpdateStream(ctx, stale) + require.NoError(t, err) + require.NoError(t, first.Close()) + + e := openEmbedded(t, dir) + require.Equal(t, 10*time.Second, streamConfig(t, e, "INGEST_acme").Duplicates, "the stale window is still on disk") + + require.NoError(t, e.SetMaxBytes(ctx, "acme", 8<<20), "same budget as before") + assert.Equal(t, EmbeddedDuplicateWindow, streamConfig(t, e, "INGEST_acme").Duplicates, + "takeStock must not have marked this pair already at budget, or this apply would have no-op'd") +} + // A durable found on disk is kept as it stands when it holds the settings // asked for — a boot over many queues writes nothing it need not — and is // updated in place when they differ; either way delivery resumes past what it From 67edd57b698590d660a24269354f865a1c5484da Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:29 -0400 Subject: [PATCH 16/22] test(mq): add an idempotency-key case to the Broker conformance suite mqtest.Run had no case for mq.WithIdempotencyKey, yet the uncertain- publish rule in the ingest handler (a claim left to lapse after a publish whose outcome is unknown, relying on the queue to drop the retry's second copy) depends on every Broker honouring it -- not just the embedded one, which already had its own duplicate-key test. Adds IdempotencyKeyDropsARepeat: a publish repeated under one key is stored once and both calls return nil; a different key is stored separately, checked via ReplaySince. Wired into the suite's case list unconditionally (no Caps flag), since every Broker must honour it. Kept internal/mq's own TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat: it additionally pins the embedded-specific mechanics of a stream opened with one duplicate window picking up EmbeddedDuplicateWindow on its next SetMaxBytes, which is outside the generic Broker contract. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/mq/mqtest/cases.go | 17 +++++++++++++++++ internal/mq/mqtest/mqtest.go | 1 + 2 files changed, 18 insertions(+) diff --git a/internal/mq/mqtest/cases.go b/internal/mq/mqtest/cases.go index 699b51945..513501de8 100644 --- a/internal/mq/mqtest/cases.go +++ b/internal/mq/mqtest/cases.go @@ -157,6 +157,23 @@ func roundTrip(t *testing.T, h Harness) { } } +// A publish repeated with the same idempotency key inside the duplicate +// window is a no-op reported as success: the ingest handler relies on this to +// make a retry of an uncertain publish (the outcome unknown after a failure +// other than a full queue) safe rather than a second copy. A different key +// is its own event. +func idempotencyKeyDropsARepeat(t *testing.T, h Harness) { + b := h.New(t) + topic := mq.Topic{Tenant: Acme, Table: "idem"} + + require.NoError(t, b.Publish(ctx(t), topic, []byte("first"), mq.WithIdempotencyKey("k1"))) + require.NoError(t, b.Publish(ctx(t), topic, []byte("repeat"), mq.WithIdempotencyKey("k1")), + "a repeat under the same key is reported as success, not stored again") + require.NoError(t, b.Publish(ctx(t), topic, []byte("second"), mq.WithIdempotencyKey("k2"))) + + replayEventually(t, b, topic, time.Time{}, []string{"first", "second"}) +} + // Nothing lands on a tenant by omission (#583), and an invalid tenant is not // backpressure a retry could clear. func refusesATopicWithoutATenant(t *testing.T, h Harness) { diff --git a/internal/mq/mqtest/mqtest.go b/internal/mq/mqtest/mqtest.go index 5978a5b6a..1d9f2edd9 100644 --- a/internal/mq/mqtest/mqtest.go +++ b/internal/mq/mqtest/mqtest.go @@ -83,6 +83,7 @@ func Run(t *testing.T, h Harness) { cases := []testCase{ {"RoundTrip", true, roundTrip}, + {"IdempotencyKeyDropsARepeat", true, idempotencyKeyDropsARepeat}, {"RefusesATopicWithoutATenant", true, refusesATopicWithoutATenant}, {"SubscribeCarriesTheTraceContext", true, subscribeCarriesTheTraceContext}, {"SubscribeSeesEveryTenant", true, subscribeSeesEveryTenant}, From a6b2108c527f1a06abcd1471b160c49f53af8f40 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:39 -0400 Subject: [PATCH 17/22] fix(app,api): correct a stale comment; don't ERROR-log a client-gone Reserve internal/app/wire.go's wirePebbleDedupe comment still described a store that fails to open on reload as a 500 "dedupe failed" -- that mapping moved to 503 "dedupe store unavailable" with Retry-After: 5 earlier in this branch's history (internal/api/reserve's dedupe.ErrUnavailable case). Update the comment to match. reserve()'s generic dd.Reserve error branch logged every failure at ERROR, including one caused by the request's own context ending (the client went away, or its deadline passed) while Reserve was in flight -- not a backend problem, and not worth paging an operator over. Check ctx.Err() and log at Debug instead when it is set; a real backend failure still logs ERROR. The response status is unchanged (moot: nothing is listening for it). TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError pins the log level via logtest; TestIngest_Dedup_ReserveError continues to pin the real-backend-failure ERROR case. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest.go | 11 ++++++++++- internal/api/ingest_test.go | 25 +++++++++++++++++++++++++ internal/app/wire.go | 5 +++-- 3 files changed, 38 insertions(+), 3 deletions(-) diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 9c92850c5..2007bb7bb 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -835,7 +835,16 @@ func (h *IngestHandler) reserve(ctx context.Context, dd dedupe.Deduplicator, tab slog.WarnContext(ctx, "dedupe store unavailable", "error", err, "table", table) return &requestAbort{Status: http.StatusServiceUnavailable, Message: "dedupe store unavailable", RetryAfter: "5"} case err != nil: - slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "table", table) + if ctx.Err() != nil { + // The request's own context ended — the client is gone, or its + // deadline passed — while Reserve was in flight. Reserve wraps + // that as an ordinary error, but it is not a backend problem + // worth an operator's attention, and the response status below + // is moot: nothing is listening for it. + slog.DebugContext(ctx, "dedupe reserve failed: request context ended", "error", err, "table", table) + } else { + slog.ErrorContext(ctx, "dedupe reserve failed", "error", err, "table", table) + } return &requestAbort{Status: http.StatusInternalServerError, Message: "dedupe failed"} } var held *dedupe.Key diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index db206fd8f..8d52815ad 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -2968,6 +2968,31 @@ func TestIngest_Dedup_ReserveError(t *testing.T) { assert.Empty(t, pub.Published()) } +// A Reserve error caused by the request's own context ending (the client +// gone, or its deadline past) is not a backend failure and must not log at +// ERROR — an operator paging on ERROR logs would otherwise be woken by +// clients that simply went away. TestIngest_Dedup_ReserveError above pins the +// real-backend-failure case, which stays ERROR. +func TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError(t *testing.T) { + buf := logtest.Capture(t, slog.LevelDebug) + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.Err = errors.New("backend down") + h := dedupHandler(t, pub, dedup, false) + + ctx, cancel := context.WithCancel(context.Background()) + cancel() + req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}).WithContext(ctx) + + w := httptest.NewRecorder() + h.Handle(w, withTenant(req)) + + assert.Contains(t, buf.String(), "dedupe reserve failed", "still logged, just not at ERROR") + assert.NotContains(t, buf.String(), `"level":"ERROR"`, "a client-gone Reserve error must not page an operator") + assert.Contains(t, buf.String(), `"level":"DEBUG"`) + assert.Empty(t, pub.Published()) +} + // #370: an explicit null id is a missing id — rejected under require_id, // published un-deduped otherwise — never the one id "" that made every // null record after the first a duplicate. diff --git a/internal/app/wire.go b/internal/app/wire.go index 20a9036fe..506b875e2 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -484,8 +484,9 @@ func (a *App) wireDedupe() error { // still closed — either the hook sees it or the boot apply reads it. An // instance that cannot open follows the registry's own rule for the shape: // flat refuses boot, like every other store, and on reload logs and leaves -// the store closed — ingest then fails closed (500 "dedupe failed") rather -// than silently publishing un-deduped, since the files asked for dedupe; +// the store closed — ingest then fails closed (503 "dedupe store +// unavailable", Retry-After: 5) rather than silently publishing un-deduped, +// since the files asked for dedupe; // nested fails closed the same way at boot too, for every tenant with // dedupe on, the next reload retrying, so it never costs the process. func (a *App) wirePebbleDedupe() error { From 5e1ff2dd1b07448b165646d6a5f211514e141fbf Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:33:30 -0400 Subject: [PATCH 18/22] test(dedupe): replace the sweep-lock timing assertion with structural ones MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits (e52b336c, this PR) raced a Commit against a sweep chunk over 300k tombstones and asserted the slowest attempt stayed under a 100ms wall-clock budget. That budget itself flaked under contention: 313ms measured when the whole internal/dedupe/... race suite ran, since test-unit runs packages in parallel under -race in make ci and a wall-clock bound moves with scheduler load. Replace it with two assertions load cannot move: a new sweepTouchHook (embedded.go, wired into deleteSweepable in sweep.go) tallies how many candidates the locked phase re-reads — asserted <= 1, the actual candidate count in this scenario, proving the locked phase's work is bounded by the candidates the unlocked read found sweepable, not by however many tombstones it silently stepped over inside Pebble to get there. A direct, non-blocking commitMu.TryLock() from within the existing sweepScanHook (fired after the unlocked read, before deleteSweepable's lock) proves the lock was free at that point without racing a goroutine or a clock at all: on the same goroutine that just did the read, TryLock fails instead of blocking if a regression left the lock held, so a pass is a direct proof, not an inference from a race won in time. Verified the new test actually catches the regression it guards against: temporarily wrapped sweepChunk's whole body (including the unlocked read) in e.commitMu.Lock()/Unlock() and dropped deleteSweepable's own lock to avoid a self-deadlock, simulating the pre-fix "sweep under the lock" design — the test failed exactly on the new TryLock assertion ("commitMu was free right after the unlocked read of 300k tombstones"), then reverted; `git diff` on sweep.go before adding the touch hook showed only the hook call, confirming a clean revert. GOTOOLCHAIN=go1.26.6 go test -race -count=20 -run Sweep ./internal/dedupe/ passed 20/20 while a concurrent `go test -race ./internal/...` ran in the background for contention (both processes exited 0). go build ./... and go vet -tags integration ./... are clean. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/embedded.go | 7 ++++- internal/dedupe/sweep.go | 3 ++ internal/dedupe/sweep_test.go | 59 ++++++++++++++++++++--------------- 3 files changed, 42 insertions(+), 27 deletions(-) diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 5eece3f3b..7cd60bd55 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -50,9 +50,14 @@ type Embedded struct { readHook func() error // sweepScanHook and sweepDeleteHook, when set, run in a sweep chunk: // between its unlocked read and its re-read, and between its re-read and - // its delete. A test races a Commit into each gap. + // its delete. A test races a Commit into each gap. sweepTouchHook, when + // set, runs once per candidate deleteSweepable re-reads under commitMu — + // a test tallies calls to pin that the locked phase's work is bounded by + // the candidate count, not by however many keys the unlocked read + // stepped over (silently, inside Pebble) to find them. sweepScanHook func() sweepDeleteHook func() + sweepTouchHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index f1c4ca180..d5a49149e 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -162,6 +162,9 @@ func (e *Embedded) deleteSweepable(db *pebble.DB, candidates [][]byte) (expired, b := db.NewBatch() defer func() { _ = b.Close() }() for _, k := range candidates { + if e.sweepTouchHook != nil { + e.sweepTouchHook() + } val, closer, err := db.Get(k) if errors.Is(err, pebble.ErrNotFound) { continue diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index 0c7ff46e9..493e37459 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -161,45 +161,52 @@ func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { // A chunk that starts over a long run of tombstones, as a tenant's version-0 // block leaves until Pebble compacts it, reads through the run without the -// lock, so a Commit racing it waits for the chunk's re-reads and deletes -// alone. Sized for the race detector, which the unit suite runs under: there, -// a chunk holding the lock across this run kept a Commit waiting ~250 ms. +// lock: deleteSweepable's locked phase only ever touches the candidates the +// unlocked read found sweepable (one here), never the tombstones stepped +// over to find them, and commitMu is free the instant that read returns. +// +// This used to race a Commit against the sweep and assert its slowest +// attempt stayed under 100ms — a wall-clock budget the race detector alone +// could push past (250ms measured once), and one that moves further under +// make ci's parallel test-unit packages (313ms measured there once, this +// PR's e52b336c: the test that introduced it flaked in its own author's +// verification). A count and a direct, non-blocking TryLock don't move with +// scheduler contention, so they replace it. func TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits(t *testing.T) { t.Parallel() e := NewEmbedded(t.TempDir()) - m := switchedOn(t, e, "acme") + switchedOn(t, e, "acme") b := e.db.NewBatch() for i := range 300_000 { require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%06d", i), nil)) } - // A version-0 key after the run, so the chunk has one to delete. + // A version-0 key after the run, so the chunk has exactly one candidate. require.NoError(t, b.Set([]byte("acme\x01"), make([]byte, 8), nil)) require.NoError(t, b.Commit(pebble.NoSync)) require.NoError(t, e.db.Flush()) - type swept struct { - res sweepResult - err error - } - done := make(chan swept, 1) - go func() { - res, err := e.sweep(context.Background(), e.db) - done <- swept{res, err} - }() - var slowest time.Duration - for i := 0; ; i++ { - select { - case s := <-done: - require.NoError(t, s.err) - require.Equal(t, sweepResult{Version0: 1}, s.res) - assert.Less(t, slowest, 100*time.Millisecond, "the slowest Commit racing the sweep") - return - default: + var touched int + e.sweepTouchHook = func() { touched++ } + var unlockedAfterScan bool + e.sweepScanHook = func() { + // Fires after the unlocked read, before deleteSweepable takes + // commitMu. TryLock needs no racing goroutine and no clock: on the + // same goroutine that just did the read, it fails instead of + // blocking if commitMu is already held — which a regression moving + // the read under the lock would leave it, self-deadlock included — + // so success here is a direct proof the read ran unlocked, not an + // inference from a race won in time. + if e.commitMu.TryLock() { + e.commitMu.Unlock() + unlockedAfterScan = true } - start := time.Now() - commitIDs(t, m, 0, fmt.Sprintf("c%d", i)) - slowest = max(slowest, time.Since(start)) } + + res, err := e.sweep(context.Background(), e.db) + require.NoError(t, err) + assert.Equal(t, sweepResult{Version0: 1}, res) + assert.True(t, unlockedAfterScan, "commitMu was free right after the unlocked read of 300k tombstones") + assert.LessOrEqual(t, touched, 1, "the locked phase touches only the sweepable candidates (1 here), not the tombstones the read stepped over to find them") } // A retention is honoured on read before any sweep has run: the key is a From 51716a29bb6b9c8f3bc7b76045a300b9e612679d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:42:28 -0400 Subject: [PATCH 19/22] test(api): pin the ERROR/DEBUG split on a Reserve failure by exact msg MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError's comment claimed TestIngest_Dedup_ReserveError pins the real-backend-failure case staying ERROR — it doesn't; that test only asserts status, body and publish count, never the log. Two mutations passed the whole internal/api suite as a result: `if ctx.Err() != nil` -> `if true` (collapses every Reserve failure onto the DEBUG branch), and the client-gone line's DebugContext -> WarnContext (both leave a bare `"level":"DEBUG"` check satisfied by an unrelated "debug: span started for ingest" line the package logs on every request). Replace the single-case test with one non-parallel, two-case TestIngest_Dedup_ReserveError_LogLevel: a live context asserts the line `"level":"ERROR","msg":"dedupe reserve failed"`; a cancelled one asserts `"level":"DEBUG","msg":"dedupe reserve failed: request context ended"`. Matching level and msg as one adjacent substring (the exact order slog's JSON handler emits them in) ties the level to the specific line rather than to the buffer as a whole, so neither mutation above can hide behind the unrelated DEBUG line. Verified both mutations now fail the new test (live-context case fails under `if true`; context-ended case fails under WarnContext), then reverted each — `git diff` on ingest.go showed no residue before the two doc fixes below were applied. Also, [MAY]: requestAbort.RetryAfter's inline comment repeated a narrower, now-stale cause list (missing the dedupe-store-unavailable 503) that duplicates and drifts from the requestAbort doc comment above it, which already owns that list — trimmed to the field's own job. mq.ErrUnavailable's doc said the API answers it with "a short Retry-After"; publishFailed sends the full dedupe lease (rounded up to whole seconds) when the failing record held a claim, and only falls back to a flat few seconds otherwise — reworded to say so. GOTOOLCHAIN=go1.26.6 go test -race ./internal/api/... ./internal/mq/... passes. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest.go | 2 +- internal/api/ingest_test.go | 66 ++++++++++++++++++++++++++----------- internal/mq/mq.go | 7 ++-- 3 files changed, 52 insertions(+), 23 deletions(-) diff --git a/internal/api/ingest.go b/internal/api/ingest.go index 2007bb7bb..c85d52c33 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -158,7 +158,7 @@ type recordReject struct { type requestAbort struct { Status int Message string - RetryAfter string // non-empty → emit a Retry-After header (503: backpressure, an unavailable broker, or an id another request holds) + RetryAfter string // non-empty → emit a Retry-After header } func (h *IngestHandler) Handle(w http.ResponseWriter, r *http.Request) { diff --git a/internal/api/ingest_test.go b/internal/api/ingest_test.go index 0e1b6d525..034a329ca 100644 --- a/internal/api/ingest_test.go +++ b/internal/api/ingest_test.go @@ -2987,29 +2987,55 @@ func TestIngest_Dedup_ReserveError(t *testing.T) { assert.Empty(t, pub.Published()) } -// A Reserve error caused by the request's own context ending (the client -// gone, or its deadline past) is not a backend failure and must not log at -// ERROR — an operator paging on ERROR logs would otherwise be woken by -// clients that simply went away. TestIngest_Dedup_ReserveError above pins the -// real-backend-failure case, which stays ERROR. -func TestIngest_Dedup_ReserveError_ContextEnded_NotLoggedAsError(t *testing.T) { - buf := logtest.Capture(t, slog.LevelDebug) - pub := &testutil.MockPublisher{} - dedup := testutil.NewMockDeduplicator() - dedup.Err = errors.New("backend down") - h := dedupHandler(t, pub, dedup, false) +// A Reserve error's log level depends on why it failed. A real backend +// failure (a live request context) stays ERROR, so an operator is paged. One +// caused by the request's own context ending (the client gone, or its +// deadline past) is not a backend problem and must log at DEBUG instead — an +// operator paging on ERROR logs would otherwise be woken by clients that +// simply went away. Not t.Parallel: it captures the process-wide default +// logger (logtest.Capture). Matched on the exact "level":"…","msg":"…" pair +// slog's JSON handler emits adjacently, not on the level alone — the package +// also logs an unrelated "debug: span started for ingest" line per request, +// which satisfies a bare `"level":"DEBUG"` check whether or not the Reserve +// line itself is DEBUG. +func TestIngest_Dedup_ReserveError_LogLevel(t *testing.T) { + tests := []struct { + name string + cancelContext bool + wantLine string + }{ + { + "live context: a real backend failure pages at ERROR", false, + `"level":"ERROR","msg":"dedupe reserve failed"`, + }, + { + "context ended: a client gone must not page, logs at DEBUG", true, + `"level":"DEBUG","msg":"dedupe reserve failed: request context ended"`, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + buf := logtest.Capture(t, slog.LevelDebug) + pub := &testutil.MockPublisher{} + dedup := testutil.NewMockDeduplicator() + dedup.Err = errors.New("backend down") + h := dedupHandler(t, pub, dedup, false) - ctx, cancel := context.WithCancel(context.Background()) - cancel() - req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}).WithContext(ctx) + ctx := context.Background() + if tt.cancelContext { + var cancel context.CancelFunc + ctx, cancel = context.WithCancel(ctx) + cancel() + } + req := ingestRequest(t, "clicks", map[string]any{"page": "/home", "event_id": "e1"}).WithContext(ctx) - w := httptest.NewRecorder() - h.Handle(w, withTenant(req)) + w := httptest.NewRecorder() + h.Handle(w, withTenant(req)) - assert.Contains(t, buf.String(), "dedupe reserve failed", "still logged, just not at ERROR") - assert.NotContains(t, buf.String(), `"level":"ERROR"`, "a client-gone Reserve error must not page an operator") - assert.Contains(t, buf.String(), `"level":"DEBUG"`) - assert.Empty(t, pub.Published()) + assert.Contains(t, buf.String(), tt.wantLine) + assert.Empty(t, pub.Published()) + }) + } } // #370: an explicit null id is a missing id — rejected under require_id, diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 6415e4812..f88bcfd54 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -192,8 +192,11 @@ var ErrQueueFull = errors.New("ingest queue is full") // ErrUnavailable is returned when the broker cannot be reached or does not // answer in time — a transient failure, not a refusal, that the API turns -// into a 503 with a short Retry-After. Only a backend whose broker is out of -// process returns it; the embedded one's publish failures are plain errors. +// into a 503. Retry-After is the dedupe lease, rounded up to whole seconds, +// when the record held a claim (so an obedient client waits out the window +// instead of retrying straight into it), else a flat few seconds. Only a +// backend whose broker is out of process returns it; the embedded one's +// publish failures are plain errors. var ErrUnavailable = errors.New("message queue unavailable") // Publisher appends events to the ingest queue. From c4147ecefefb0686057b16ab2a20a39c85b1a538 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:53:41 -0400 Subject: [PATCH 20/22] test(dedupe): pin sweepCandidates' read as unlocked from inside its loop MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits claimed more than it checked. sweepTouchHook fired once per element of `candidates`, and the fixture has exactly one visible key, so `touched <= 1` could never fail: a full keyspace walk added inside the locked phase would still pass (R4). The TryLock ran in sweepScanHook, which fires AFTER sweepCandidates returns, so a regression that read under commitMu and released the lock just before that hook still passed (R3) — only "the whole chunk under one lock" was actually caught (R5). The 300k-tombstone fixture changed no verdict either way (a handful of tombstones behaves identically) while costing 1.65s alone / 3.1s under package load in a 15s-timeout package, and the comment recorded this test's own history (citing a commit that a squash would erase) instead of what it asserts. Drop sweepTouchHook (field, wiring, assertion). To pin "the read runs unlocked" for real, sweepCandidates now takes an onKey hook called once per key from inside its own loop, before evaluating it — wired through sweepChunk as e.sweepReadHook. Renamed TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits to TestEmbedded_SweepReadRunsUnlocked: TryLock/Unlock from inside that hook, while the read is still running, so a lock held anywhere during the read is caught in the act rather than inferred from whether it was released before some later checkpoint. The fixture shrinks to a handful of tombstones ahead of the one live key, since the verdict never depended on the count. Comment cut to the one thing the test asserts; the old wall-clock/hook history belongs here instead. Verified by mutation, each applied then reverted (`git diff internal/dedupe/sweep.go` clean before the real change was made): - R3 (only the read under the lock, released right after): wrapped just the sweepCandidates call in sweepChunk with commitMu.Lock()/Unlock() — TestEmbedded_SweepReadRunsUnlocked failed ("commitMu must be free while sweepCandidates' read is running"). - R5 (the whole chunk under one lock): wrapped sweepChunk's body in commitMu.Lock()/defer Unlock() and dropped deleteSweepable's own lock to avoid a self-deadlock — ran ONLY the target test (not the package: TestEmbedded_SweepNeverDeletesACommitLandingMidChunk self-deadlocks under this mutation, racing a Commit against a sweep that never releases the lock) — failed with the same assertion. GOTOOLCHAIN=go1.26.6 go test -race -count=5 -run Sweep ./internal/dedupe/ passes (5/5, including TestEmbedded_SweepReadRunsUnlocked and every other Sweep-prefixed test); the full package (go test -race -count=1 ./internal/dedupe/...) passes too, now in ~5s rather than the prior fixture's ~57s at -count=20. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/embedded.go | 11 ++++---- internal/dedupe/sweep.go | 13 +++++----- internal/dedupe/sweep_test.go | 47 ++++++++++++----------------------- 3 files changed, 28 insertions(+), 43 deletions(-) diff --git a/internal/dedupe/embedded.go b/internal/dedupe/embedded.go index 7cd60bd55..227564b71 100644 --- a/internal/dedupe/embedded.go +++ b/internal/dedupe/embedded.go @@ -50,14 +50,13 @@ type Embedded struct { readHook func() error // sweepScanHook and sweepDeleteHook, when set, run in a sweep chunk: // between its unlocked read and its re-read, and between its re-read and - // its delete. A test races a Commit into each gap. sweepTouchHook, when - // set, runs once per candidate deleteSweepable re-reads under commitMu — - // a test tallies calls to pin that the locked phase's work is bounded by - // the candidate count, not by however many keys the unlocked read - // stepped over (silently, inside Pebble) to find them. + // its delete. A test races a Commit into each gap. sweepReadHook, when + // set, runs once per key sweepCandidates' unlocked read visits, from + // inside its loop — a test TryLocks commitMu there to prove the read + // itself never holds it, not just the instant after it returns. sweepScanHook func() sweepDeleteHook func() - sweepTouchHook func() + sweepReadHook func() } // NewEmbedded returns the embedded implementation under dataDir. Nothing is diff --git a/internal/dedupe/sweep.go b/internal/dedupe/sweep.go index d5a49149e..25759f39b 100644 --- a/internal/dedupe/sweep.go +++ b/internal/dedupe/sweep.go @@ -104,7 +104,7 @@ func (e *Embedded) sweep(ctx context.Context, db *pebble.DB) (sweepResult, error // read and a run of them left by an earlier pass would otherwise hold every // Commit for its whole length. func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, res *sweepResult) ([]byte, error) { - candidates, next, err := sweepCandidates(db, from, e.now()) + candidates, next, err := sweepCandidates(db, from, e.now(), e.sweepReadHook) if err != nil || len(candidates) == 0 { return next, err } @@ -128,14 +128,18 @@ func (e *Embedded) sweepChunk(ctx context.Context, db *pebble.DB, from []byte, r // sweepCandidates reads the next sweepChunk keys, starting at from, and // returns those sweepable at now and where the next chunk starts (nil at the -// end). -func sweepCandidates(db *pebble.DB, from []byte, now time.Time) (candidates [][]byte, next []byte, err error) { +// end). onKey, when non-nil, runs once per key visited, before it is +// evaluated — a test hook proving this read holds no lock while it runs. +func sweepCandidates(db *pebble.DB, from []byte, now time.Time, onKey func()) (candidates [][]byte, next []byte, err error) { it, err := db.NewIter(&pebble.IterOptions{LowerBound: from}) if err != nil { return nil, nil, fmt.Errorf("dedupe sweep: %w", err) } seen := 0 for valid := it.First(); valid; valid = it.Next() { + if onKey != nil { + onKey() + } if seen == sweepChunk { next = bytes.Clone(it.Key()) break @@ -162,9 +166,6 @@ func (e *Embedded) deleteSweepable(db *pebble.DB, candidates [][]byte) (expired, b := db.NewBatch() defer func() { _ = b.Close() }() for _, k := range candidates { - if e.sweepTouchHook != nil { - e.sweepTouchHook() - } val, closer, err := db.Get(k) if errors.Is(err, pebble.ErrNotFound) { continue diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index 493e37459..a0255e398 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -159,54 +159,39 @@ func TestEmbedded_SweepNeverDeletesACommitLandingMidChunk(t *testing.T) { assert.True(t, dup, "the commit made mid-chunk survived the sweep") } -// A chunk that starts over a long run of tombstones, as a tenant's version-0 -// block leaves until Pebble compacts it, reads through the run without the -// lock: deleteSweepable's locked phase only ever touches the candidates the -// unlocked read found sweepable (one here), never the tombstones stepped -// over to find them, and commitMu is free the instant that read returns. -// -// This used to race a Commit against the sweep and assert its slowest -// attempt stayed under 100ms — a wall-clock budget the race detector alone -// could push past (250ms measured once), and one that moves further under -// make ci's parallel test-unit packages (313ms measured there once, this -// PR's e52b336c: the test that introduced it flaked in its own author's -// verification). A count and a direct, non-blocking TryLock don't move with -// scheduler contention, so they replace it. -func TestEmbedded_SweepChunkOverTombstonesDoesNotHoldCommits(t *testing.T) { +// sweepCandidates' read never holds commitMu, over a fixture with a few +// tombstones ahead of the one live key it finds sweepable. +func TestEmbedded_SweepReadRunsUnlocked(t *testing.T) { t.Parallel() e := NewEmbedded(t.TempDir()) switchedOn(t, e, "acme") b := e.db.NewBatch() - for i := range 300_000 { - require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%06d", i), nil)) + for i := range 4 { + require.NoError(t, b.Delete(fmt.Appendf(nil, "acme\x00%02d", i), nil)) } - // A version-0 key after the run, so the chunk has exactly one candidate. require.NoError(t, b.Set([]byte("acme\x01"), make([]byte, 8), nil)) require.NoError(t, b.Commit(pebble.NoSync)) require.NoError(t, e.db.Flush()) - var touched int - e.sweepTouchHook = func() { touched++ } - var unlockedAfterScan bool - e.sweepScanHook = func() { - // Fires after the unlocked read, before deleteSweepable takes - // commitMu. TryLock needs no racing goroutine and no clock: on the - // same goroutine that just did the read, it fails instead of - // blocking if commitMu is already held — which a regression moving - // the read under the lock would leave it, self-deadlock included — - // so success here is a direct proof the read ran unlocked, not an - // inference from a race won in time. + var sawLocked bool + e.sweepReadHook = func() { + // TryLock from inside the still-running read needs no racing + // goroutine and no clock: on the same goroutine doing the read, it + // fails instead of blocking if commitMu is already held, so a + // regression that reads under the lock is caught while the read is + // still in progress — not inferred from a race won in time, and not + // missable by a lock released just before some later checkpoint. if e.commitMu.TryLock() { e.commitMu.Unlock() - unlockedAfterScan = true + } else { + sawLocked = true } } res, err := e.sweep(context.Background(), e.db) require.NoError(t, err) assert.Equal(t, sweepResult{Version0: 1}, res) - assert.True(t, unlockedAfterScan, "commitMu was free right after the unlocked read of 300k tombstones") - assert.LessOrEqual(t, touched, 1, "the locked phase touches only the sweepable candidates (1 here), not the tombstones the read stepped over to find them") + assert.False(t, sawLocked, "commitMu must be free while sweepCandidates' read is running") } // A retention is honoured on read before any sweep has run: the key is a From 21380681f230f02e72f13b6d44d67c080e46a784 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:55:04 -0400 Subject: [PATCH 21/22] test: keep this branch's two broker stores in storedir TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat and the api package's realPipeline opened the embedded broker on a bare t.TempDir(). Both replay through a disk-backed consumer, so they were exposed to the late consumer-state write (#442) that storedir absorbs. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/api/ingest_window_test.go | 3 ++- internal/mq/embedded_test.go | 2 +- 2 files changed, 3 insertions(+), 2 deletions(-) diff --git a/internal/api/ingest_window_test.go b/internal/api/ingest_window_test.go index ba350eedd..73bb83bb8 100644 --- a/internal/api/ingest_window_test.go +++ b/internal/api/ingest_window_test.go @@ -15,6 +15,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/settings" "github.com/Wave-RF/WaveHouse/internal/testutil" + "github.com/Wave-RF/WaveHouse/internal/testutil/storedir" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" ) @@ -289,7 +290,7 @@ func (p *faultyPublisher) Publish(ctx context.Context, topic mq.Topic, data []by // in the tenant's queue. func realPipeline(t *testing.T, fail func(call int) (bool, error)) (*IngestHandler, func() int) { t.Helper() - broker, err := mq.NewEmbedded(t.TempDir()) + broker, err := mq.NewEmbedded(storedir.New(t)) require.NoError(t, err) t.Cleanup(func() { _ = broker.Close() }) require.NoError(t, broker.SetMaxBytes(t.Context(), testStore.Tenant(), 64<<20)) diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index a2fced640..e1d28281e 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -151,7 +151,7 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // path, where takeStock itself must not mistake a stale window for one // already at budget. func TestEmbeddedNATS_Publish_IdempotencyKeyDropsARepeat(t *testing.T) { - e := openEmbedded(t, t.TempDir()) + e := openEmbedded(t, storedir.New(t)) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() // Explicit rather than the server's default, which happens to match today. From a76bbb9d86a6188b07413b9c6e64afe9d51b85e1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:00:32 -0400 Subject: [PATCH 22/22] test(dedupe): assert the sweep's read hook ran TestEmbedded_SweepReadRunsUnlocked asserted only that the hook never saw commitMu held, which also holds if the hook never fires: passing nil for the hook left it green. It now counts the visits and requires one; with the hook unwired it fails. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/dedupe/sweep_test.go | 10 ++++------ 1 file changed, 4 insertions(+), 6 deletions(-) diff --git a/internal/dedupe/sweep_test.go b/internal/dedupe/sweep_test.go index a0255e398..ee87088e2 100644 --- a/internal/dedupe/sweep_test.go +++ b/internal/dedupe/sweep_test.go @@ -173,14 +173,11 @@ func TestEmbedded_SweepReadRunsUnlocked(t *testing.T) { require.NoError(t, b.Commit(pebble.NoSync)) require.NoError(t, e.db.Flush()) + var visits int var sawLocked bool e.sweepReadHook = func() { - // TryLock from inside the still-running read needs no racing - // goroutine and no clock: on the same goroutine doing the read, it - // fails instead of blocking if commitMu is already held, so a - // regression that reads under the lock is caught while the read is - // still in progress — not inferred from a race won in time, and not - // missable by a lock released just before some later checkpoint. + visits++ + // Non-blocking, on the reading goroutine: fails if the read holds commitMu. if e.commitMu.TryLock() { e.commitMu.Unlock() } else { @@ -191,6 +188,7 @@ func TestEmbedded_SweepReadRunsUnlocked(t *testing.T) { res, err := e.sweep(context.Background(), e.db) require.NoError(t, err) assert.Equal(t, sweepResult{Version0: 1}, res) + assert.Positive(t, visits, "the read hook ran") assert.False(t, sawLocked, "commitMu must be free while sweepCandidates' read is running") }