From 67df3173993f30d18cf92c00810fa1e886346ad6 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:56:52 -0400 Subject: [PATCH 1/8] perf(cache): flat version index, pruned per tenant The local version index nested each table under its tenant's version and each scope under its table's, and was never pruned: every InvalidateTenant left the tenant's whole index behind. It now holds one version per tenant, (tenant, table) and (tenant, table, scope), bumped in place. A tenant bump drops the tenant's index and its next key gets a process-unique generation; a table bump drops the table's scopes. After each reload the wiring prunes the index to the tenants served. Fixes #262 for the local backend. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 1 + docs/src/content/docs/architecture.md | 2 +- internal/app/app_test.go | 42 +++++ internal/app/wire.go | 16 +- internal/cache/local.go | 19 +- internal/cache/local_test.go | 33 ++++ internal/cache/version_manager.go | 232 +++++++++++++++++-------- internal/cache/version_manager_test.go | 192 ++++++++++++++------ 9 files changed, 406 insertions(+), 135 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index d3e34545..6207f60a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,9 +29,9 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `LocalCache.Prune` drops its cache version index), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `...
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 83f2fd83..50ea8f2e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,6 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) bounds its versions with a TTL instead. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3ba9a9eb..de4f6dfc 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -111,7 +111,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d9..1f665b6a 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -16,6 +16,7 @@ import ( "os" "path/filepath" "strings" + "sync" "sync/atomic" "syscall" "testing" @@ -635,6 +636,47 @@ func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { assert.Equal(t, []tenant.ID{"globex", "acme"}, mock.GetTenants(), "restored: the same") } +// pruneRecorder is a cache that records, at each Prune, which of the tenants +// it is asked about are still served. +type pruneRecorder struct { + testutil.MockCache + mu sync.Mutex + served []map[tenant.ID]bool +} + +func (p *pruneRecorder) Prune(served func(tenant.ID) bool) { + p.mu.Lock() + defer p.mu.Unlock() + p.served = append(p.served, map[tenant.ID]bool{"acme": served("acme"), "globex": served("globex")}) +} + +func (p *pruneRecorder) last() map[tenant.ID]bool { + p.mu.Lock() + defer p.mu.Unlock() + if len(p.served) == 0 { + return nil + } + return p.served[len(p.served)-1] +} + +// Every reload prunes the cache's version index down to the tenants served, +// so a tenant rejected or removed stops holding it (#262). +func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) + a := newApp(t, testConfig(t, root), Options{}) + rec := &pruneRecorder{} + a.cache = rec + + rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) + a.tenants.Reload("test") + assert.Equal(t, map[tenant.ID]bool{"acme": true, "globex": false}, rec.last(), "rejected") + + rewriteSettings(t, filepath.Join(root, "globex"), nil) + require.NoError(t, os.RemoveAll(filepath.Join(root, "acme"))) + a.tenants.Reload("test") + assert.Equal(t, map[tenant.ID]bool{"acme": false, "globex": true}, rec.last(), "removed; the repaired one served again") +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} diff --git a/internal/app/wire.go b/internal/app/wire.go index 4c9f1da2..8811ba63 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -591,7 +591,16 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// pruner is a cache whose version index lives in the process and would +// otherwise keep a tenant that stopped being served (cache.LocalCache). +type pruner interface { + Prune(served func(tenant.ID) bool) +} + +// wireCache opens the L1 cache — the only tier in standalone mode. After +// every reload a tenant no longer served, removed or rejected alike, has its +// version index dropped (#262); its cache is orphaned with it, as it would +// be anyway when it came back (wireClickHouse). func (a *App) wireCache() error { l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) if err != nil { @@ -600,6 +609,11 @@ func (a *App) wireCache() error { // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc a.cache = l1 a.add(component{name: "cache", close: withoutContext(l1.Close)}) + a.tenants.AfterAdopt(func([]tenant.ID) { + if p, ok := a.cache.(pruner); ok { + p.Prune(a.served) + } + }) return nil } diff --git a/internal/cache/local.go b/internal/cache/local.go index 1a755e05..3e0e14b5 100644 --- a/internal/cache/local.go +++ b/internal/cache/local.go @@ -69,10 +69,11 @@ func (l *LocalCache) Set(_ context.Context, snap Snapshot, value []byte, ttl tim // view. Returns the number of namespaces processed. // // This bumps exactly what it's given. A whole-table bump already subsumes every -// per-scope bump for the same table (the table version is embedded in every -// namespace key), so a caller that knows a whole-table bump is coming should drop -// the now-redundant scope entries itself — the ingest worker does this as it -// builds the batch, where it already loops once and knows it's a single table. +// per-scope bump for the same table (every key that folds a scope version +// folds the table version too), so a caller that knows a whole-table bump is +// coming should drop the now-redundant scope entries itself — the ingest +// worker does this as it builds the batch, where it already loops once and +// knows it's a single table. func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint64, error) { for _, ns := range namespaces { if ns.Scope == "" { @@ -85,13 +86,21 @@ func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint } // InvalidateTenant orphans every cached result of tenant id, pipe results -// included: one version bump, nothing enumerated (see +// included: its version index is dropped, nothing enumerated (see // VersionManager.BumpTenant). func (l *LocalCache) InvalidateTenant(_ context.Context, id tenant.ID) error { l.versionManager.BumpTenant(id) return nil } +// Prune drops the version index of every tenant served rejects, orphaning +// its entries as InvalidateTenant would, so a tenant removed or rejected at +// a reload stops holding memory (#262). The entries themselves go with +// their TTL or Ristretto's eviction. +func (l *LocalCache) Prune(served func(tenant.ID) bool) { + l.versionManager.Prune(served) +} + // Wait blocks until all buffered writes have been applied. // Exposed for testing; production callers rarely need this. func (l *LocalCache) Wait() { diff --git a/internal/cache/local_test.go b/internal/cache/local_test.go index 2fdb3237..ad77a40d 100644 --- a/internal/cache/local_test.go +++ b/internal/cache/local_test.go @@ -2,10 +2,13 @@ package cache_test import ( "testing" + "time" + "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil/cachetest" ) @@ -23,3 +26,33 @@ func TestLocalCache_Conformance(t *testing.T) { t.Parallel() cachetest.Run(t, newLocal, cachetest.Options{MaxValueBytes: localMaxCost}) } + +// A tenant that stops being served has its index dropped: what it cached is +// orphaned — it misses when served again — and a tenant still served keeps +// its entries. +func TestLocalCache_Prune(t *testing.T) { + t.Parallel() + c, err := cache.NewLocal(localMaxCost) + require.NoError(t, err) + t.Cleanup(func() { _ = c.Close() }) + ctx := t.Context() + fill := func(id tenant.ID) { + _, snap, err := c.Lookup(ctx, id, "q", []cache.Namespace{{Tenant: id, Table: "events"}}) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) + } + get := func(id tenant.ID) []byte { + e, _, err := c.Lookup(ctx, id, "q", []cache.Namespace{{Tenant: id, Table: "events"}}) + require.NoError(t, err) + return e.Value + } + fill("acme") + fill("globex") + c.Wait() + require.NotNil(t, get("acme")) + require.NotNil(t, get("globex")) + + c.Prune(func(id tenant.ID) bool { return id == "acme" }) + assert.NotNil(t, get("acme"), "still served") + assert.Nil(t, get("globex"), "pruned: orphaned, never revived") +} diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d97bc70a..dfe18b8c 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -9,27 +9,44 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" ) -// VersionManager handles the safe tracking of table + scope versioning. -// It uses a standard map because versions must NEVER be evicted under memory pressure. -// TODO: this potentially could be bad/dangerous with a low amount of RAM available/high memory pressure AND a TON of tables/scopes per table... will need to work out eventually +// VersionManager is the invalidation index: one version per tenant, per +// (tenant, table) and per (tenant, table, scope), in maps keyed by name +// alone, never by another version (#262). A bump overwrites a version in +// place, so the index holds one entry per live tenant, table and scope +// however often each is bumped, and forgetting a tenant releases all of it. +// +// A query key folds all three versions of each dependency, which gives the +// lattice: a table bump orphans every scope, a scope bump that scope and the +// whole-table view, and a tenant bump everything of the tenant's. type VersionManager struct { mu sync.RWMutex - // tenantVersions leads every key of a tenant, so BumpTenant orphans the - // tenant's every namespace and query in one step — the ones no bump ever - // keyed included, which is what an enumeration of the maps would miss. - tenantVersions map[tenant.ID]uint64 // -> tenant_version - tableVersions map[string]uint64 // ..
-> table_version - namespaceVersions map[string]uint64 // ..
.. -> namespace_version + // tenants holds each tenant's index from the first query key built for + // it until the tenant is bumped or pruned. + tenants map[tenant.ID]*tenantVersions + + // lastGen is the last generation handed to a tenant; see tenantVersions.gen. + lastGen uint64 } -// NewVersionManager initializes the thread-safe version store. -func NewVersionManager() *VersionManager { - return &VersionManager{ - tenantVersions: make(map[tenant.ID]uint64), - tableVersions: make(map[string]uint64), - namespaceVersions: make(map[string]uint64), - } +// tenantVersions is one tenant's slice of the index. +type tenantVersions struct { + // gen is the tenant's version: unique within the process, so a tenant + // forgotten and recreated can never fold a generation an entry was + // cached under. That is what makes dropping the tenant's whole index a + // safe bump. + gen uint64 + tables map[string]*tableVersions +} + +// tableVersions is one table's version and its scopes'. A missing table or +// scope reads as 0: an entry is only ever removed together with a bump of +// the version above it (a table bump clears the scopes, a tenant bump +// drops the tables), so a 0 read after a removal never matches an entry +// cached before it. +type tableVersions struct { + version uint64 + scopes map[string]uint64 } // Namespace is one (tenant, table, scope) a cached result depends on. The @@ -41,76 +58,101 @@ type Namespace struct { Scope string } -// tableKeyLocked renders the table-versions key, -// "..
"; caller must hold vm.mu. A tenant id -// cannot contain a dot and callers encode the table dot-free, so the tokens -// can never run together. -func (vm *VersionManager) tableKeyLocked(id tenant.ID, table string) string { - return fmt.Sprintf("%s.%d.%s", id, vm.tenantVersions[id], table) -} - -// namespaceKeyLocked builds the namespace-table key; caller must hold vm.mu. -func (vm *VersionManager) namespaceKeyLocked(ns Namespace) string { - tk := vm.tableKeyLocked(ns.Tenant, ns.Table) - return fmt.Sprintf("%s.%d.%s", tk, vm.tableVersions[tk], ns.Scope) -} - -// NamespaceKey renders the namespace-table key for ns at its tenant's and -// table's current versions: -// "..
.." (scopeless -// scope is "", so e.g. ".0.
.."). -func (vm *VersionManager) NamespaceKey(ns Namespace) string { - vm.mu.RLock() - defer vm.mu.RUnlock() - return vm.namespaceKeyLocked(ns) +// NewVersionManager initializes the thread-safe version store. +func NewVersionManager() *VersionManager { + return &VersionManager{tenants: make(map[tenant.ID]*tenantVersions)} } // QueryKey builds the queries-table key for tenant id's result that depends // on deps: the query's sha (hash of SQL+params) folded with the tenant's -// version and every dependency's namespace key AND its namespace version, so -// a bump of the tenant or of any dependency misses the key — a result with no -// deps (a pipe) is orphaned by BumpTenant too. A structured query passes one -// Namespace; a pipe passes several. Deps are sorted so their order never -// changes the key. +// version and, for every dependency, its tenant's, table's and scope's +// versions, so a bump of the tenant or of any dependency misses the key — a +// result with no deps (a pipe) is orphaned by BumpTenant too. A structured +// query passes one Namespace, a pipe none yet (#343). Deps are sorted so +// their order never changes the key. Every version is read under one lock, +// so the key is one consistent snapshot. +// +// The first key built for a tenant creates its index at a fresh generation. func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) string { + vm.mu.RLock() + key, ok := vm.queryKeyLocked(id, sha, deps, false) + vm.mu.RUnlock() + if ok { + return key + } + vm.mu.Lock() + defer vm.mu.Unlock() + key, _ = vm.queryKeyLocked(id, sha, deps, true) + return key +} + +// queryKeyLocked renders QueryKey with vm.mu held — for writing when create +// is set, which creates the index of each tenant the key names that has +// none; otherwise such a tenant reports !ok. +func (vm *VersionManager) queryKeyLocked(id tenant.ID, sha string, deps []Namespace, create bool) (string, bool) { + index := func(id tenant.ID) (*tenantVersions, bool) { + tv := vm.tenants[id] + if tv == nil && create { + tv = vm.newTenantLocked(id) + } + return tv, tv != nil + } + own, ok := index(id) + if !ok { + return "", false + } segs := make([]string, len(deps)) - // Lock per dependency rather than across the whole loop: each dep's table + - // namespace versions are read together (consistent for that dep), but we don't - // hold the lock across all deps. A concurrent bump can land between deps; the - // caller files its fill under this key (a Snapshot), so a bump that lands - // anywhere after the read of a version orphans it. The sort/join run with - // no lock held. for i, d := range deps { - vm.mu.RLock() - nsKey := vm.namespaceKeyLocked(d) - segs[i] = fmt.Sprintf("%s.%d", nsKey, vm.namespaceVersions[nsKey]) - vm.mu.RUnlock() + tv, ok := index(d.Tenant) + if !ok { + return "", false + } + var table, scope uint64 + if t := tv.tables[d.Table]; t != nil { + table, scope = t.version, t.scopes[d.Scope] + } + segs[i] = fmt.Sprintf("%s.%d.%s.%d.%s.%d", d.Tenant, tv.gen, d.Table, table, d.Scope, scope) } - vm.mu.RLock() - tv := vm.tenantVersions[id] - vm.mu.RUnlock() sort.Strings(segs) - return fmt.Sprintf("%s|%s.%d|%s", sha, id, tv, strings.Join(segs, "|")) + return fmt.Sprintf("%s|%s.%d|%s", sha, id, own.gen, strings.Join(segs, "|")), true +} + +func (vm *VersionManager) newTenantLocked(id tenant.ID) *tenantVersions { + vm.lastGen++ + tv := &tenantVersions{gen: vm.lastGen, tables: make(map[string]*tableVersions)} + vm.tenants[id] = tv + return tv +} + +// tableLocked is the entry for a tenant's table, created at version 0, or +// nil when the tenant has no index: no key folds its current generation +// yet, so there is nothing a bump could orphan. Caller holds vm.mu for +// writing. +func (vm *VersionManager) tableLocked(id tenant.ID, table string) *tableVersions { + tv := vm.tenants[id] + if tv == nil { + return nil + } + t := tv.tables[table] + if t == nil { + t = &tableVersions{} + tv.tables[table] = t + } + return t } // BumpTable advances a tenant's table version, orphaning every namespace — and // every cached query — that depends on the table, in one step (the whole-table -// nuke). The same table under another tenant is untouched. +// nuke). The table's scope versions are dropped with it: every key they were +// folded into also folds the old table version. The same table under another +// tenant is untouched. func (vm *VersionManager) BumpTable(id tenant.ID, table string) { vm.mu.Lock() defer vm.mu.Unlock() - vm.tableVersions[vm.tableKeyLocked(id, table)]++ -} - -// BumpTenant advances a tenant's version, orphaning its every namespace — -// and every cached query, whatever its deps — in one step (the whole-tenant -// nuke): every namespace and query key of the tenant carries the version, so -// nothing has to be enumerated, and a table no bump ever keyed is orphaned -// like the rest. Other tenants are untouched. -func (vm *VersionManager) BumpTenant(id tenant.ID) { - vm.mu.Lock() - defer vm.mu.Unlock() - vm.tenantVersions[id]++ + if t := vm.tableLocked(id, table); t != nil { + t.version++ + t.scopes = nil + } } // BumpNamespace advances one (tenant, table, scope) namespace plus the table's @@ -119,8 +161,54 @@ func (vm *VersionManager) BumpTenant(id tenant.ID) { func (vm *VersionManager) BumpNamespace(ns Namespace) { vm.mu.Lock() defer vm.mu.Unlock() - vm.namespaceVersions[vm.namespaceKeyLocked(ns)]++ + t := vm.tableLocked(ns.Tenant, ns.Table) + if t == nil { + return + } + if t.scopes == nil { + t.scopes = make(map[string]uint64) + } + t.scopes[ns.Scope]++ if ns.Scope != "" { - vm.namespaceVersions[vm.namespaceKeyLocked(Namespace{Tenant: ns.Tenant, Table: ns.Table})]++ + t.scopes[""]++ + } +} + +// BumpTenant orphans every cached query of a tenant, whatever its deps, in +// one step (the whole-tenant nuke), by dropping the tenant's index: the next +// key built for it gets a fresh generation, which no cached entry folds. +// Nothing has to be enumerated, a table no bump ever keyed is orphaned like +// the rest, and the index the tenant held is released. Other tenants are +// untouched. +func (vm *VersionManager) BumpTenant(id tenant.ID) { + vm.mu.Lock() + defer vm.mu.Unlock() + delete(vm.tenants, id) +} + +// Prune drops the index of every tenant keep rejects, as BumpTenant would, +// so a tenant that stops being served stops holding memory; one served again +// starts over at a fresh generation. +func (vm *VersionManager) Prune(keep func(tenant.ID) bool) { + vm.mu.Lock() + defer vm.mu.Unlock() + for id := range vm.tenants { + if !keep(id) { + delete(vm.tenants, id) + } + } +} + +// size is the number of versions the index holds, for tests. +func (vm *VersionManager) size() int { + vm.mu.RLock() + defer vm.mu.RUnlock() + n := len(vm.tenants) + for _, tv := range vm.tenants { + n += len(tv.tables) + for _, t := range tv.tables { + n += len(t.scopes) + } } + return n } diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index e1a1c330..c9278577 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -8,45 +8,17 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" ) -func TestVersionManager_NamespaceKey(t *testing.T) { - t.Parallel() - vm := NewVersionManager() - - // The tenant leads at its default version (0), then the table at its - // default version (0); a scopeless namespace renders a trailing dot. The - // flat directory's tenant is "0". - assert.Equal(t, "acme.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.0.users.0.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) - assert.Equal(t, "0.0.users.0.", vm.NamespaceKey(Namespace{Tenant: tenant.Default, Table: "users"})) - - // The table version is embedded in every namespace key for that tenant's - // table, so a BumpTable is reflected across all its scopes at once — and - // nowhere else: the same table under another tenant keeps its version. - vm.BumpTable("acme", "users") - assert.Equal(t, "acme.0.users.1.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.0.users.1.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) - assert.Equal(t, "globex.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "globex", Table: "users"})) - - // The tenant version leads every key of the tenant, so a BumpTenant moves - // every table of acme's — the never-bumped orders table included — to a - // fresh key space, at table version 0 again, and no other tenant's. - vm.BumpTenant("acme") - assert.Equal(t, "acme.1.users.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.1.orders.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "orders"})) - assert.Equal(t, "globex.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "globex", Table: "users"})) -} - func TestVersionManager_QueryKey(t *testing.T) { t.Parallel() vm := NewVersionManager() - // One dependency at default versions: - // sha | . | ..
.... + // sha | . | ..
...; + // acme's index is created by its first key, at generation 1. key := vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}}) - assert.Equal(t, "hash123|acme.0|acme.0.users.0.org_1.0", key) + assert.Equal(t, "hash123|acme.1|acme.1.users.0.org_1.0", key) // No deps (a pipe) still folds the tenant version. - assert.Equal(t, "hash123|acme.0|", vm.QueryKey("acme", "hash123", nil)) + assert.Equal(t, "hash123|acme.1|", vm.QueryKey("acme", "hash123", nil)) // Dependency order must not change the key (segments are sorted). deps1 := []Namespace{{Tenant: "acme", Table: "a"}, {Tenant: "acme", Table: "b"}} @@ -59,6 +31,9 @@ func TestVersionManager_QueryKey(t *testing.T) { vm.QueryKey("acme", "h", []Namespace{{Tenant: "acme", Table: "users"}}), vm.QueryKey("globex", "h", []Namespace{{Tenant: "globex", Table: "users"}})) assert.NotEqual(t, vm.QueryKey("acme", "h", nil), vm.QueryKey("globex", "h", nil)) + + // Reading keys is stable: nothing but a bump moves a version. + assert.Equal(t, key, vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}})) } func TestVersionManager_BumpTable(t *testing.T) { @@ -69,16 +44,16 @@ func TestVersionManager_BumpTable(t *testing.T) { orders := []Namespace{{Tenant: "acme", Table: "orders", Scope: "org_1"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - usersBefore := vm.QueryKey(users[0].Tenant, "h", users) - ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) - globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + usersBefore := vm.QueryKey("acme", "h", users) + ordersBefore := vm.QueryKey("acme", "h", orders) + globexBefore := vm.QueryKey("globex", "h", globexUsers) // Bumping a table changes the key for that tenant's table but leaves other // tables — and the same table under another tenant — alone. vm.BumpTable("acme", "users") - assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) - assert.Equal(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders)) - assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey("acme", "h", users)) + assert.Equal(t, ordersBefore, vm.QueryKey("acme", "h", orders)) + assert.Equal(t, globexBefore, vm.QueryKey("globex", "h", globexUsers)) } func TestVersionManager_BumpNamespace(t *testing.T) { @@ -90,24 +65,44 @@ func TestVersionManager_BumpNamespace(t *testing.T) { otherScope := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_2"}} otherTenant := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - scopedBefore := vm.QueryKey(scoped[0].Tenant, "h", scoped) - wholeBefore := vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable) - otherBefore := vm.QueryKey(otherScope[0].Tenant, "h", otherScope) - otherTenantBefore := vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant) + scopedBefore := vm.QueryKey("acme", "h", scoped) + wholeBefore := vm.QueryKey("acme", "h", wholeTable) + otherBefore := vm.QueryKey("acme", "h", otherScope) + otherTenantBefore := vm.QueryKey("globex", "h", otherTenant) // Bumping (acme, users, org_1) changes that scope AND the whole-table view, // but leaves every other scope — and the same scope under another tenant — // valid. vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) - assert.NotEqual(t, scopedBefore, vm.QueryKey(scoped[0].Tenant, "h", scoped)) - assert.NotEqual(t, wholeBefore, vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable)) - assert.Equal(t, otherBefore, vm.QueryKey(otherScope[0].Tenant, "h", otherScope)) - assert.Equal(t, otherTenantBefore, vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant)) + assert.NotEqual(t, scopedBefore, vm.QueryKey("acme", "h", scoped)) + assert.NotEqual(t, wholeBefore, vm.QueryKey("acme", "h", wholeTable)) + assert.Equal(t, otherBefore, vm.QueryKey("acme", "h", otherScope)) + assert.Equal(t, otherTenantBefore, vm.QueryKey("globex", "h", otherTenant)) +} + +// A table bump drops the table's scope versions, which then read as 0 again +// — safe only because every key a scope version was folded into also folds +// the table version the bump moved. Pinned so a table bump that forgot to +// advance the table version would revive the scoped entry. +func TestVersionManager_BumpTableDropsScopes(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + scoped := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}} + + fresh := vm.QueryKey("acme", "h", scoped) + vm.BumpNamespace(scoped[0]) + bumped := vm.QueryKey("acme", "h", scoped) + vm.BumpTable("acme", "users") + after := vm.QueryKey("acme", "h", scoped) + + assert.NotEqual(t, fresh, after) + assert.NotEqual(t, bumped, after) + assert.Equal(t, 2, vm.size(), "the tenant and its table; the scopes went with the table bump") } // TestVersionManager_BumpTenant: a tenant's every namespace is orphaned in -// one step — a table that was never bumped (so has no key of its own to bump) -// included — and no other tenant's is touched. +// one step — a table that was never bumped included — and no other tenant's +// is touched. func TestVersionManager_BumpTenant(t *testing.T) { t.Parallel() vm := NewVersionManager() @@ -115,18 +110,107 @@ func TestVersionManager_BumpTenant(t *testing.T) { users := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}} orders := []Namespace{{Tenant: "acme", Table: "orders"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} + vm.QueryKey("acme", "h", users) vm.BumpTable("acme", "users") - usersBefore := vm.QueryKey(users[0].Tenant, "h", users) - ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) - globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + usersBefore := vm.QueryKey("acme", "h", users) + ordersBefore := vm.QueryKey("acme", "h", orders) + globexBefore := vm.QueryKey("globex", "h", globexUsers) vm.BumpTenant("acme") - assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) - assert.NotEqual(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders), "a table no bump ever keyed is orphaned too") - assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey("acme", "h", users)) + assert.NotEqual(t, ordersBefore, vm.QueryKey("acme", "h", orders), "a table no bump ever keyed is orphaned too") + assert.Equal(t, globexBefore, vm.QueryKey("globex", "h", globexUsers)) pipeBefore := vm.QueryKey("acme", "h", nil) vm.BumpTenant("acme") assert.NotEqual(t, pipeBefore, vm.QueryKey("acme", "h", nil), "a result with no deps is orphaned too") } + +// Dropping a tenant's index is a bump only because the index it gets back +// never repeats a generation: every key built before any of these drops must +// differ from every key built after it. A counter per tenant restarting at 0 +// fails this, reviving the first entry. +func TestVersionManager_GenerationsNeverRepeat(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + deps := []Namespace{{Tenant: "acme", Table: "users"}} + seen := map[string]bool{} + for i := range 100 { + key := vm.QueryKey("acme", "h", deps) + assert.False(t, seen[key], "round %d revived %s", i, key) + seen[key] = true + if i%2 == 0 { + vm.BumpTenant("acme") + } else { + vm.Prune(func(tenant.ID) bool { return false }) + } + } +} + +// A bump of a tenant with no index is a no-op: no key folds its next +// generation yet, so nothing needs orphaning — and an insert still in flight +// for a tenant just pruned does not bring its index back. +func TestVersionManager_BumpWithoutIndex(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + vm.BumpTable("acme", "users") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + vm.BumpTenant("acme") + assert.Zero(t, vm.size()) +} + +func TestVersionManager_Prune(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + acme := []Namespace{{Tenant: "acme", Table: "users"}} + globex := []Namespace{{Tenant: "globex", Table: "users"}} + acmeBefore := vm.QueryKey("acme", "h", acme) + globexBefore := vm.QueryKey("globex", "h", globex) + vm.BumpTable("acme", "users") + vm.BumpTable("globex", "users") + acmeBumped := vm.QueryKey("acme", "h", acme) + globexBumped := vm.QueryKey("globex", "h", globex) + + vm.Prune(func(id tenant.ID) bool { return id == "globex" }) + assert.Equal(t, 2, vm.size(), "globex and its table; acme released") + assert.Equal(t, globexBumped, vm.QueryKey("globex", "h", globex), "a kept tenant is untouched") + + back := vm.QueryKey("acme", "h", acme) + assert.NotEqual(t, acmeBefore, back, "a pruned tenant never revives what it cached") + assert.NotEqual(t, acmeBumped, back) + assert.NotEqual(t, globexBefore, globexBumped) +} + +// The index holds one version per live tenant, table and scope, however often +// each is bumped (#262): the nested index this replaced kept every table and +// scope under every tenant version it had seen. +func TestVersionManager_SizeDoesNotGrowWithBumps(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + touch := func() { + for _, id := range []tenant.ID{"acme", "globex"} { + for _, table := range []string{"users", "orders"} { + for _, scope := range []string{"", "org_1", "org_2"} { + vm.QueryKey(id, "h", []Namespace{{Tenant: id, Table: table, Scope: scope}}) + vm.BumpNamespace(Namespace{Tenant: id, Table: table, Scope: scope}) + } + } + } + } + touch() + settled := vm.size() + assert.Equal(t, 2+2*2+2*2*3, settled, "two tenants, two tables each, three scopes each") + + for i := range 10_000 { + switch i % 3 { + case 0: + vm.BumpTable("acme", []string{"users", "orders"}[i%2]) + case 1: + vm.BumpTenant("globex") + } + touch() + assert.LessOrEqual(t, vm.size(), settled) + } + assert.Equal(t, settled, vm.size()) +} From ed8022bf34127e5cdd751d916f7d7dde162971fc Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:06:14 -0400 Subject: [PATCH 2/8] docs(cache): name the cache prune hook; pin LocalCache as a pruner Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/app/wire.go | 9 +++++++-- 3 files changed, 9 insertions(+), 4 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 6207f60a..adbec1d8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `LocalCache.Prune` drops its cache version index), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index de4f6dfc..368c9368 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/internal/app/wire.go b/internal/app/wire.go index 8811ba63..900c1999 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -143,8 +143,9 @@ func gapWindows(tenants *settings.Registry) map[tenant.ID]time.Duration { } // served reports whether the registry is serving tenant id: what the -// per-tenant resources — verifiers, dedupe stores, open streams — are pruned -// by once a reload removes or rejects their tenant. +// per-tenant resources — verifiers, dedupe stores, open streams, the cache +// version index — are pruned by once a reload removes or rejects their +// tenant. func (a *App) served(id tenant.ID) bool { _, ok := a.tenants.For(id) return ok @@ -597,6 +598,10 @@ type pruner interface { Prune(served func(tenant.ID) bool) } +// The hook below asserts pruner at run time; this keeps LocalCache from +// silently dropping out of it. +var _ pruner = (*cache.LocalCache)(nil) + // wireCache opens the L1 cache — the only tier in standalone mode. After // every reload a tenant no longer served, removed or rejected alike, has its // version index dropped (#262); its cache is orphaned with it, as it would From ee51e8401cac9e4e8669a6924dcd6b1931e5445a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:50:24 -0400 Subject: [PATCH 3/8] test(app): assert the wired cache is a pruner before swapping it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestReload_PrunesCacheIndexToServedTenants replaced a.cache with a recorder by hand, so the run-time a.cache.(pruner) check against the cache wireCache actually builds was never exercised: a change that stopped wiring a pruning cache left the test green while pruning silently went dead. Assert the wired cache satisfies pruner before the swap. Also reword a stale assertion message ("a rejection releases; it orphans nothing yet") that no longer describes real wiring now that a rejection drops the tenant's cache index via Prune, and correct the CHANGELOG's claim that the Redis-compatible backend (#613) already bounds its versions with a TTL — that backend does not exist yet. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- internal/app/app_test.go | 5 ++++- 2 files changed, 5 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2515caa5..674c2df0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. -- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) bounds its versions with a TTL instead. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 84c67070..d21f376a 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -671,7 +671,7 @@ func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) a.tenants.Reload("test") - assert.Empty(t, mock.GetTenants(), "a rejection releases; it orphans nothing yet") + assert.Empty(t, mock.GetTenants(), "a rejection calls no InvalidateTenant (Prune drops its index)") rewriteSettings(t, filepath.Join(root, "globex"), nil) _, adopted = a.tenants.Reload("test") require.True(t, adopted) @@ -712,6 +712,9 @@ func (p *pruneRecorder) last() map[tenant.ID]bool { func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) a := newApp(t, testConfig(t, root), Options{}) + _, ok := a.cache.(pruner) + require.True(t, ok, "the wired cache prunes") + rec := &pruneRecorder{} a.cache = rec From 4df73f995fd14139088e310b2175faf27eebc69d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:46:54 -0400 Subject: [PATCH 4/8] docs(cache): clarify which bumps a no-index tenant absorbs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The version_manager.go bullet read "a bump of a tenant with no index is a no-op" right after introducing BumpTenant, so it looked scoped to that one call. It actually covers any bump — table, scope or tenant — against a tenant with no index, which is what lets Prune and a departed tenant's index stay gone (neither the sharedTables fan-out nor an insert still in flight for a just-pruned tenant can revive it). Found by pre-push review of #621 (origin/feat/cache-snapshot...HEAD). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/architecture.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4690761d..7aede5ba 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration From 95575f27a49982cddd1c1febd72028ff669a901b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:54:37 -0400 Subject: [PATCH 5/8] docs(cache): CHANGELOG says #262 is part-fixed, not fixed The #262 entry said "fixes #262 for the in-process cache", which would close the issue on merge (delete_branch_on_merge + squash_merge_commit_message: PR_BODY carry the PR body's own "Fixes #262" into main's history the same way). #262 has a residual left open on purpose: per-table scope cardinality isn't capped, since scope stays empty until #235 populates it. Reworded to "part of #262" and named the residual, matching the PR body and the issue comment that records the trigger. Found by pre-push review of #621 (origin/feat/cache-snapshot...HEAD). Co-Authored-By: Claude Sonnet 5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a89a0d2b..6631b3c4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -89,7 +89,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. -- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): part of [#262](https://github.com/Wave-RF/WaveHouse/issues/262) (growth across bumps and a departed tenant's memory; per-table scope cardinality is left open, see the issue) and of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. From 558bb23835691a2ca43c4b3b7afb85ce3d755972 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:30:43 -0400 Subject: [PATCH 6/8] fix(cache): pin the no-index bump test and fix a "query key" mixup MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestVersionManager_BumpWithoutIndex asserted only on vm.size() after its final BumpTenant call, which unconditionally deletes the tenant's index regardless of what the earlier no-index bumps did — so it could not fail against a tableLocked that wrongly creates an index instead of returning nil. Assert after each bump instead, and add the pruned-tenant case a write racing Prune must not revive. Also reword two docs passages that overloaded or misstated a term: architecture.md used "query key" for both the caller's input and the rendered entry key in the same paragraph; AGENTS.md's cache bullet read as if the index maps were keyed by escaped names; CHANGELOG.md's #382 bullet said the version index builds its keys with internal/keyenc, contradicting its own closing sentence that the index's entries are unaffected by the escaping. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/cache/version_manager_test.go | 18 +++++++++++++++++- 4 files changed, 20 insertions(+), 4 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a504ffd1..a74dc6bb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), each raw table and scope name escaped by `keyenc` where the key is built (a `Namespace` carries them raw, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight, escaped whole as the lead field of the stored key `|.||…`, where each raw table and scope name is escaped by `keyenc` (a `Namespace` carries them raw, so no caller escapes); the index holds a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by raw name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 6631b3c4..14de673f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -88,7 +88,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the cache renders each result's key with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): part of [#262](https://github.com/Wave-RF/WaveHouse/issues/262) (growth across bumps and a departed tenant's memory; per-table scope cardinality is left open, see the issue) and of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7aede5ba..678c52a3 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). An entry's key (`QueryKey`) folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index 84792054..07945306 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -164,10 +164,26 @@ func TestVersionManager_GenerationsNeverRepeat(t *testing.T) { func TestVersionManager_BumpWithoutIndex(t *testing.T) { t.Parallel() vm := NewVersionManager() + vm.BumpTable("acme", "users") + assert.Zero(t, vm.size(), "a table bump for a tenant with no index creates nothing") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + assert.Zero(t, vm.size(), "a namespace bump for a tenant with no index creates nothing") + vm.BumpTenant("acme") - assert.Zero(t, vm.size()) + assert.Zero(t, vm.size(), "bumping a tenant with no index is a no-op") + + // An insert still in flight for a tenant just pruned must not bring its + // index back: a write racing the prune sees the tenant gone and bumps + // blind, same as above. + vm.QueryKey("acme", "h", nil) + vm.Prune(func(tenant.ID) bool { return false }) + assert.Zero(t, vm.size(), "prune released the tenant's index") + + vm.BumpTable("acme", "users") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + assert.Zero(t, vm.size(), "a bump for a tenant just pruned must not recreate its index") } func TestVersionManager_Prune(t *testing.T) { From 568f2ddd20a826be68fd7bc8800e2fd2d7524df5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:46:46 -0400 Subject: [PATCH 7/8] docs(cache): version_manager.go's own doc carries the query-key/entry-key mixup too Same overload as architecture.md's version_manager.go bullet: the type doc's "A query key folds all three versions of each dependency" means the rendered entry key QueryKey returns, not the caller's input. Reword to match. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/version_manager.go | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d67a6077..d6fd7496 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -16,9 +16,10 @@ import ( // place, so the index holds one entry per live tenant, table and scope // however often each is bumped, and forgetting a tenant releases all of it. // -// A query key folds all three versions of each dependency, which gives the -// lattice: a table bump orphans every scope, a scope bump that scope and the -// whole-table view, and a tenant bump everything of the tenant's. +// An entry's key (`QueryKey`) folds all three versions of each dependency, +// which gives the lattice: a table bump orphans every scope, a scope bump +// that scope and the whole-table view, and a tenant bump everything of the +// tenant's. type VersionManager struct { mu sync.RWMutex From c42b6db2603dc8d517441bba981dc6685833c135 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:04:44 -0400 Subject: [PATCH 8/8] docs(cache): fix the one query-key/entry-key spot the terminology pass missed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit version_manager.go's tenants field comment still said "the first query key built for it" — the same overload the type doc three lines above and the other doc fixes in this round eliminated (query key = the caller's input, entry key = what QueryKey renders). Both confirmation reviewers caught this independently. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/version_manager.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d6fd7496..6b1de3b5 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -23,8 +23,8 @@ import ( type VersionManager struct { mu sync.RWMutex - // tenants holds each tenant's index from the first query key built for - // it until the tenant is bumped or pruned. + // tenants holds each tenant's index from the first entry key (`QueryKey`) + // built for it until the tenant is bumped or pruned. tenants map[tenant.ID]*tenantVersions // lastGen is the last generation handed to a tenant; see tenantVersions.gen.