Area: mq / cache / dedupe / ingest / app — distributed deployment epic · scoped with @EricAndrechek , 2026-09-24
#583 makes one process serve many tenants. This epic makes many processes serve them: every stateful layer that is embedded today gets a shared, external implementation, the background loops can run outside the API process, and a deployment chooses the implementation of each layer at boot. The embedded implementations stay the default, so a standalone operator notices no change. The work is being built alongside #583 and ships after it.
Decisions (2026-09-24)
Implementations are chosen per process, not per tenant. The boot config (config.yaml / WH_*) picks each layer's implementation; a tenant's config.json keeps only its own tunables (budgets, TTLs, enabled switches).
External NATS: WaveHouse does not own the streams. A NATS Kubernetes operator creates and owns the streams, consumers and policies, shared across tenants rather than one set per tenant. WaveHouse checks at boot that they exist and never creates or edits them. The embedded implementation keeps its per-tenant streams (feat(mq): give every tenant a queue of its own #612 ).
Remote dedupe stays pluggable. Whether a tenant is pinned to a home region or must be deduplicated globally is undecided. DynamoDB in one region is built first; the multi-region-strong-consistency and Scylla paths are designed and documented.
Each workstream is a draft PR, stacked where one depends on another. Nothing merges before epic(settings): multi-tenant settings directory and tenant-scoped runtime #583 is done.
Workstreams
PRs (all merged to main by 2026-09-30)
Every stack was squash-merged to main, children collapsed into their roots: #647 , #632 , #616 , #655 , #619 (with #627 ), #618 , #615 , #622 , #623 , #614 (E stack), #625 (F stack), and #624 (the whole D stack plus B2: #636 , #639 , #644 , #646 , and the fix rounds #695 /#696 ).
A: fix(ingest): retry ClickHouse outages instead of dead-lettering #619 ClickHouse outage is retried, never dead-lettered → fix(api): map ClickHouse query failures by class, not HTTP status #627 query paths map CH errors by class (fixes query API maps ClickHouse caller-side errors (e.g. ACCESS_DENIED) to HTTP 502 instead of 4xx #403 , bug(api): request-caused query errors return HTTP 500 + retryable:true (should be 4xx, non-retryable) #271 )
G/B/C: feat(config): choose each layer's implementation at boot #618 G1 choose each layer's implementation at boot → feat(coord): leases, in-process implementation #615 B1 internal/coord leases, sweeper elected → feat(app): process roles #622 C1 process roles + ops-only listener
C3: fix(discovery): jitter the refresh retry backoff #616 discovery retry jitter (closes discovery: add jitter to RetryRefresh exponential backoff for clustered mode #141 ), on main
E: feat(cache): shared redis cache, snapshot keys, uncached write pipes #614 E1 snapshot at lookup, tenant token on every key (fixes Cache: result can be keyed to a newer version than the data it reflects (write-during-query staleness) #382 ) → perf(cache): flat version index, pruned per tenant #621 E2 flat local version index (perf(cache): cache version-manager map grows unbounded with tables × scopes #262 ) · feat(cache): shared cache backend on Redis, Valkey and Dragonfly #626 E3 Redis/Valkey/Dragonfly backend → feat(config): cache.backend=redis selects the shared cache #630 E4 cache.backend: redis wiring + two-instance test · fix(pipes): run write pipes every call, uncached and uncoalesced #634 write pipes uncached (fixes bug(pipes): mutation pipe results are cached and coalesced — the write silently drops on repeat calls #386 ), on E1
F: fix(dedupe)!: reserve/commit ids, windowed ingest, retention, dynamodb #625 F1 Reserve/Commit/Release, tenant+table keys, atomic Pebble (fixes bug(dedupe): CheckAndMark is not atomic — concurrent same-id requests both pass #390 , Per-table dedupe id_field (+ fix cross-table dedupe keyspace collision) #222 , bug(ingest): explicit null id_field slips past require_id and dedupes every row onto "<nil>" #370 ) → fix(ingest): windowed reserve/publish/commit; 503 when dedupe down #629 F2 windowed ingest + Nats-Msg-Id (fixes bug(ingest): dedupe marks the event id before the NATS publish — a failed publish + client retry permanently drops the event #384 ) → feat(dedupe): retention per tenant and table, and an expiry sweep #633 F4 retention + Pebble sweep (fixes Optional TTL/size bound on the dedupe store + durability docs #220 ) · feat(dedupe): add a DynamoDB backend, tested but not yet wired #628 F3 DynamoDB backend → feat(app): choose the DynamoDB dedupe backend at boot #635 F5 dedupe.backend: dynamodb wiring
D: test(mq): one conformance suite for every Broker #623 D1 mqtest conformance + ErrUnavailable · feat(mq): external nats broker with sharded, pinned ingest workers #624 D2 operator topology, verifier, manifests (S1 passed) → feat(mq): a Broker over an external NATS cluster #636 D3 ExternalNATS broker → feat(app): mq.backend selects embedded or external NATS #639 D4 mq.backend: nats wiring → fix(mq): drain partitions a lower N leaves; quiet close #644 close-log and partition-shrink fixes → feat(mq): coord leases on a NATS KV bucket #646 B2 NATS KV leases (carries test(mq,app): fit the unit budget; fail fast on an uncreatable store #647 's commits until test(mq,app): fit the unit budget; fail fast on an uncreatable store #647 lands)
Found along the way: test(app): internal/app unit tests use 8–17 s of the 15 s budget, time out under load #617 slow internal/app/internal/mq unit tests → test(mq,app): fit the unit budget; fail fast on an uncreatable store #647 speed-up (on main, ~5 s per package) · fix(api): timeouts and memory limits on uncapped query paths are retried #620 · bug(config): an explicit false/0 in config.yaml is replaced by the field's env-default #631 YAML zero replaced by default → fix(config): keep an explicit false/0/"" from config.yaml #632 fix, on main
Integration proof (closed, superseded): test(e2e): api, ingest and sweeper in separate processes #645 merged every stack, plus B2, the test(app): internal/app unit tests use 8–17 s of the 15 s budget, time out under load #617 test speed-up and C2 (api, ingest and sweeper as separate processes over NATS, Redis and DynamoDB), onto main. make ci passes there. It is not meant to merge directly; its body lists which integration fix goes back to which PR.
Related: #583 , #612 , #246 , #247 , #264 , #384 , #390 , #393 , #403 , #271 , #220 , #222 .
Status (2026-09-30)
Done: every workstream is on main; the last, D with B2, landed in #624 . The embedded backends remain the default. Before production traffic relies on external NATS, the 3-node checks in #694 should run.
Follow-ups, each tracked on its own:
mq / NATS: bug(mq): mixed-version N/V rollout can double-pull shards; version floor only sees the connected server #684 , bug(mq): history replay can skip rows after a reset, or return early during a leader election #685 , bug(mq): RESET can redeliver rows a live, draining process still holds #686 , bug(mq): storage-limit errors (10002/10023) map to 500 instead of 503 + Retry-After #687 , bug(mq): closing while a Consume is in progress reports a spurious failure #688 , mq: durable/partition topology checks miss backoff, pause_until, max_waiting, subject_transform, sources #689 , bug(mq): DuplicateWindow() returns the shortest window; retention warning needs the longest #690 , bug(config): nats mq defaults duplicate mq.Default* untested, no grammar check, sweeper role is a silent no-op #691 , bug(mq): internal/mq integration tests are selected by name pattern, so a new prefix is silently skipped #692 , mq: follow-ups from #624 — membership cost, Unowned batching, crash-restart wait, allocations, docs, per-tenant budgets #693 , mq(nats): pre-production measurements on a 3-node cluster #694 , docs(mq): NATS sizing advice, error code and ack wording are wrong after #624 #697
ingest: ingest: a hot table is capped at one worker, and a dead worker's partitions wait out the lease #649 , ingest: accept events durably while the dedupe backend is down #652 , mq: one unacked event pins the embedded ingest stream's purge for every table of its tenant #653 , ingest: split oversized batches by halving, and don't split TOO_MANY_PARTS caused by merge backpressure #657 , bug(ingest): a stop during a slow claim tick can outrun ShutdownTimeout and reorder rows #698 , bug(ingest): a stale unbound timestamp makes a later bind failure log ERROR at once #699 , test(ingest): TestRejectPoison_CountedByDisposition fails at -count>1 #700
dedupe: dedupe: cross-region dedupe on DynamoDB MRSC needs a lapsed-claim sweeper #651 · api: api: ClickHouse ACCESS_DENIED on /v1/query and pipes should be 502, not 403 #650 · cache: bug(cache): a demoted Redis primary reads as healthy; Sentinel mode is unconfigured #656 , bug(cache): Redis recovery probe must fit a fresh dial in the op timeout #664 , bug(cache): standalone mode against a cluster-enabled Redis misses silently #669
Area: mq / cache / dedupe / ingest / app — distributed deployment epic · scoped with @EricAndrechek, 2026-09-24
#583 makes one process serve many tenants. This epic makes many processes serve them: every stateful layer that is embedded today gets a shared, external implementation, the background loops can run outside the API process, and a deployment chooses the implementation of each layer at boot. The embedded implementations stay the default, so a standalone operator notices no change. The work is being built alongside #583 and ships after it.
Decisions (2026-09-24)
config.yaml/WH_*) picks each layer's implementation; a tenant'sconfig.jsonkeeps only its own tunables (budgets, TTLs, enabled switches).Workstreams
mq.Broker. Shared streams owned by the operator, built on feat(mq): give every tenant a queue of its own #612's per-tenant interfaces.cache.Cachewith version namespaces shared across pods, so an invalidation on one pod reaches every pod.internal/app.PRs (all merged to main by 2026-09-30)
Every stack was squash-merged to
main, children collapsed into their roots: #647, #632, #616, #655, #619 (with #627), #618, #615, #622, #623, #614 (E stack), #625 (F stack), and #624 (the whole D stack plus B2: #636, #639, #644, #646, and the fix rounds #695/#696).internal/coordleases, sweeper elected → feat(app): process roles #622 C1 process roles + ops-only listenerRetryRefreshexponential backoff for clustered mode #141), on maincache.backend: rediswiring + two-instance test · fix(pipes): run write pipes every call, uncached and uncoalesced #634 write pipes uncached (fixes bug(pipes): mutation pipe results are cached and coalesced — the write silently drops on repeat calls #386), on E1dedupe.backend: dynamodbwiringmqtestconformance +ErrUnavailable· feat(mq): external nats broker with sharded, pinned ingest workers #624 D2 operator topology, verifier, manifests (S1 passed) → feat(mq): a Broker over an external NATS cluster #636 D3ExternalNATSbroker → feat(app): mq.backend selects embedded or external NATS #639 D4mq.backend: natswiring → fix(mq): drain partitions a lower N leaves; quiet close #644 close-log and partition-shrink fixes → feat(mq): coord leases on a NATS KV bucket #646 B2 NATS KV leases (carries test(mq,app): fit the unit budget; fail fast on an uncreatable store #647's commits until test(mq,app): fit the unit budget; fail fast on an uncreatable store #647 lands)internal/app/internal/mqunit tests → test(mq,app): fit the unit budget; fail fast on an uncreatable store #647 speed-up (on main, ~5 s per package) · fix(api): timeouts and memory limits on uncapped query paths are retried #620 · bug(config): an explicit false/0 in config.yaml is replaced by the field's env-default #631 YAML zero replaced by default → fix(config): keep an explicit false/0/"" from config.yaml #632 fix, on mainmake cipasses there. It is not meant to merge directly; its body lists which integration fix goes back to which PR.Related: #583, #612, #246, #247, #264, #384, #390, #393, #403, #271, #220, #222.
Status (2026-09-30)
Done: every workstream is on
main; the last, D with B2, landed in #624. The embedded backends remain the default. Before production traffic relies on external NATS, the 3-node checks in #694 should run.Follow-ups, each tracked on its own: