Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
63 commits
Select commit Hold shift + click to select a range
a890511
chore: pin the core dependency to a tag instead of tracking main (#25)
fylorn Sep 23, 2026
a90e9f3
feat!: forward what can be forwarded, convert only what must be (#26)
fylorn Sep 23, 2026
ca7647c
refactor: take the circuit-breaker registry and metric labels back fr…
fylorn Sep 23, 2026
a6678cf
refactor: redact PII with core's guard engine, and restore tool argum…
fylorn Sep 23, 2026
437a671
feat: inspect the tool calls an upstream returns (#29)
fylorn Sep 23, 2026
038796b
build: keep only line tables in dev and test builds (#30)
fylorn Sep 24, 2026
05f52e8
feat: one circuit-breaker state machine, and hidden characters in req…
fylorn Sep 24, 2026
8a76395
fix: make the integration suite pass again (#32)
fylorn Sep 24, 2026
6482690
test: fix the two flaky integration tests (#33)
fylorn Sep 24, 2026
9f46116
fix: bill a Chat stream whose caller did not ask for usage (#34)
fylorn Sep 24, 2026
3c35ab8
fix: a request the upstream refuses no longer fails over or trips bre…
fylorn Sep 24, 2026
234fcce
Merge main (v1.1.0) into dev
fylorn Sep 24, 2026
82f8061
docs: cut releases from a release branch, and keep the tag's headings…
fylorn Sep 24, 2026
bd57b5d
ci: run checks on pull requests into dev, including the integration s…
fylorn Sep 24, 2026
c1b2085
fix: bill cached input at cache prices, estimate missing usage, strip…
fylorn Sep 24, 2026
ad5b75b
refactor: take back what only this side used from core (#41)
fylorn Sep 24, 2026
3a617f3
feat(gateway): take official hosts, output fallback and error shapes …
fylorn Sep 24, 2026
ed1c8b5
feat(gateway): accept Gemini clients, and the Responses API over a We…
fylorn Sep 24, 2026
28d4387
test: count the probes that miss before the streaming cache hit (#44)
fylorn Sep 24, 2026
d3f36fc
refactor(server): move the catalog's SQL into repositories (#45)
fylorn Sep 24, 2026
83e9bcc
refactor(server): move the dashboard, limits and log-forwarding handl…
fylorn Sep 24, 2026
9b5b326
refactor(server): move the MCP handlers' SQL into repositories (#50)
fylorn Sep 24, 2026
86beb82
refactor(gateway): run the request guards on core's tw-guard engines …
fylorn Sep 24, 2026
1c80b83
refactor(server): move the identity handlers' SQL into repositories (…
fylorn Sep 24, 2026
c25b44f
fix(mcp): revoking a default connection no longer fails with a 500 (#51)
fylorn Sep 24, 2026
dbce845
test: leave the cancelled stream after it has started, not on a timer…
fylorn Sep 24, 2026
c3d2d57
refactor(server): move the access handlers' SQL into repositories (#48)
fylorn Sep 24, 2026
d92c2d1
fix(settings): read security.totp_required as a boolean, and seed aut…
fylorn Sep 24, 2026
7e95183
feat(gateway): record a client that leaves before its response exists…
fylorn Sep 24, 2026
367df3c
feat(auth): enforce the TOTP requirement (#55)
fylorn Sep 24, 2026
443de8d
feat(gateway): keep the last response on a Responses WebSocket (#56)
fylorn Sep 24, 2026
6890c72
Merge main (v2.0.0) into dev
fylorn Sep 24, 2026
d4f6d5a
docs(release): CI runs the integration suite on the release PR (#58)
fylorn Sep 24, 2026
ef3de87
docs(contributing): point vulnerability reports at the organization's…
fylorn Sep 25, 2026
c4f8b1c
docs(readme): current description of ThinkWatch Lite and ThinkWatch C…
fylorn Sep 25, 2026
f259962
feat(providers): authenticate Bedrock with an API key (#61)
fylorn Sep 28, 2026
78f9ab5
test: hold the early-cancel client once its key is known, not on a ti…
fylorn Sep 28, 2026
1b752f2
feat(providers): list Bedrock's models, and let the route editor take…
fylorn Sep 28, 2026
1f08421
fix(bedrock): refuse keys that will not decrypt, and keep instance-ro…
fylorn Sep 28, 2026
e480bfd
refactor(bedrock): use core's shared tw-bedrock, and core v0.55.0 (#65)
fylorn Sep 29, 2026
e1e7999
docs(readme): rewrite both READMEs to be short and accurate (#66)
fylorn Sep 30, 2026
af9eb81
Merge main (v2.1.0) into dev
fylorn Sep 30, 2026
ed6dd39
docs(readme): 37 MCP templates, what the setup wizard does, body reda…
fylorn Sep 30, 2026
9afcc08
fix(rbac): make the seeded team_manager work at team scope (#69)
fylorn Sep 30, 2026
fa3a5b9
fix(rbac): require gateway use, and count only grants that give it (#70)
fylorn Sep 30, 2026
a1eed6f
Merge main (v2.2.0) into dev
fylorn Sep 30, 2026
da88b09
Merge main (sponsor line) into dev
fylorn Oct 2, 2026
1a399e7
Allow clippy::double_must_use on the async_trait BlobStore (#73)
fylorn Oct 2, 2026
8d8d96f
Unify the request guards with thinkwatch-core's rule model (#72)
fylorn Oct 3, 2026
2baa123
Fix what the pre-release review of the guard unification found (#75)
fylorn Oct 3, 2026
33c22dc
Merge main (v3.0.0) into dev
fylorn Oct 3, 2026
dcc8616
Merge main (README: ThinkWatch Lite for individual developers) into dev
fylorn Oct 4, 2026
6b73344
Merge main (v3.1.0) into dev
fylorn Oct 5, 2026
03c393c
Usage limits: token limits refuse, key limits count on the key, Retry…
fylorn Oct 5, 2026
4075855
fix(redis): connect over TLS for rediss:// URLs (#82)
fylorn Oct 5, 2026
33234ee
Concurrent startup, body retention on restart, and Helm network polic…
fylorn Oct 5, 2026
a7c1d49
fix(redis): keep deleting past a failed SCAN or DEL in pattern deletes
fylorn Oct 5, 2026
c9d0127
docs(limits): name sliding::admit and sliding::record in comments
fylorn Oct 5, 2026
a4b442a
docs: don't advertise ClickHouse over HTTPS
fylorn Oct 5, 2026
fead598
fix(startup): give the startup probe room for queued schema setups
fylorn Oct 5, 2026
325ceb0
fix(helm): de-duplicate Sentinel ports, require redis.caSecret.key
fylorn Oct 5, 2026
6ba5ea1
docs(changelog): complete what to read before upgrading to 3.2.0
fylorn Oct 5, 2026
9588c4a
chore(release): tag 3.2.0
fylorn Oct 5, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -26,11 +26,23 @@ DB_MAX_CONNECTIONS=10
DB_PASSWORD=__SECRET_HEX_16__
# dev: DATABASE_URL=postgres://${DB_USER}:${DB_PASSWORD}@localhost:5432/${DB_NAME}
# prod: DATABASE_URL=postgres://${DB_USER}:${DB_PASSWORD}@postgres:5432/${DB_NAME}?sslmode=disable
# Postgres directly, or a pooler in session mode: never one in transaction
# mode (PgBouncer pool_mode = transaction). Start-up schema setup holds a
# session-level advisory lock (deploy/helm/think-watch/README.md).

# --- Redis ---
REDIS_PASSWORD=__SECRET_HEX_16__
# dev: REDIS_URL=redis://:${REDIS_PASSWORD}@localhost:6379
# prod: REDIS_URL=redis://:${REDIS_PASSWORD}@redis:6379
# Redis Cluster: REDIS_URL=redis-cluster://:${REDIS_PASSWORD}@redis-0:6379?node=redis-1:6379
# (see deploy/helm/think-watch/README.md, "Redis Cluster")
# TLS (managed Redis services usually require it): the rediss:// scheme,
# rediss-cluster:// for a cluster. The certificate is checked against the
# public CAs and the host name in the URL.
# REDIS_URL=rediss://:${REDIS_PASSWORD}@my-cache.example.com:6379
# A Redis whose certificate a private CA signed: a PEM file with that CA,
# then trusted alone (see deploy/helm/think-watch/README.md, "Redis over TLS")
# REDIS_CA_CERT=/etc/thinkwatch/redis-ca.crt

# --- Application ---
JWT_SECRET=__SECRET_HEX_32__
Expand Down
154 changes: 153 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,157 @@ target.

## [Unreleased]

## [3.2.0] — 2026-10-05

Redis Cluster and Redis over TLS now work, so managed Redis services can be
used. Usage limits work as documented: token limits refuse, an API key's
limits count on that key, and a refused request counts against nothing.
Server instances can start together against one database, a restart no
longer clears captured bodies older than 30 days, and the Helm network
policy allows the ports the database URLs name. The thinkwatch-core crates
stay at v0.62.0.

### Read before upgrading

- **Rate-limit windows start empty.** Rate-limit counters move to new Redis
keys (one hash per counter, tagged so that Redis Cluster can run them), and
the counts from before the upgrade are not carried over: every window starts
empty and fills from the first request after the upgrade. The old keys
expire by themselves within two window lengths. Users' budget counters are
kept; API keys' are not (see below).
- **Route health starts fresh.** A route's samples, circuit breaker and
lifetime request count move to new keys for the same reason, so every route
starts closed with nothing counted. The old lifetime counters never expire;
`redis-cli --scan --pattern 'route_health:[0-9a-f]*' | xargs redis-cli del`
removes them (the new keys start `route_health:{`).
- **An API key's own budgets start again.** 3.1.0 counted a key's budget on
its owner's counter (`budget:user:<user id>:…`); 3.2.0 counts it on the
key's own (`budget:api_key_lineage:<lineage id>:…`), which starts at zero.
A key with a monthly budget of its own can spend all of it again in the
rest of the month.
- **Limits are looser while the rollout runs.** Pods of 3.1.0 and 3.2.0
count rate limits and route health on different keys, so each sees only
its own version's requests: limits let more through, and circuit breakers
can disagree, until the last 3.1.0 pod is gone.
- **Token limits start refusing.** A `tokens` rate limit never refused a
request before. It now does once its window is full, so a deployment with
token limits will see `429`s where it saw none.
- **A key's limits no longer replace its owner's.** A rate limit or budget set
on an API key used to take the place of the owner's limit for the same
window or period. Both now apply, each on its own counter: a key's limits can
narrow what its owner may do through that key, never widen it. A key given a
higher limit than its owner to give it more room needs the owner's limit
raised instead.
- **`rediss://` connects over TLS.** 3.1.0 connected to a `rediss://` URL
over plain TCP. A `rediss://` URL pointing at a port without TLS now fails
at start: point it at the TLS port, or write `redis://`.
- The certificate must name the host in a subjectAltName. One that names
it only in its CN, which `redis-cli` accepts, is refused.
- Cluster nodes that announce IP addresses their certificates don't name
can't be reached.
- **`budget_unavailable` in dashboards.** With Redis down and
`security.rate_limit_fail_closed` on, a request that has a budget is
refused as `budget_unavailable`, not `rate_limiter_unavailable`: budgets
are checked first.
- **Upgrade with an ordinary rollout.** Earlier versions don't take the lock
that now makes instances set up the schema one at a time, so don't restart
instances of the old version while the first one of this version starts.
- **No transaction-mode pooler in front of Postgres.** Schema setup now holds a
Postgres session-level advisory lock, which a pooler in transaction mode
(PgBouncer `pool_mode = transaction`) can leave held, and every later start
then waits for it: point `DATABASE_URL` at Postgres itself or at a pooler in
session mode (the Helm chart's README has the details).

### Fixed

- **Token limits refuse requests.** A `tokens` rate limit never refused
anything, and stopped counting once a request would have taken it past its
limit. A request is now refused once the window's recorded usage reaches the
limit, and every request's tokens are recorded after it, even past the limit
— a window can overshoot by what was in flight when it filled.
- **Several request limits at once.** With two or more `requests` limits on a
user (per minute and per hour, say), every request that passed was counted
twice, and only one of the limits could refuse; with
`security.rate_limit_fail_closed` on, every request was refused as
`rate_limiter_unavailable`.
- **An API key's limits count on that key.** They were counted on its owner's
counter, which every key of the owner shared, and the usage the console
reads for a key (`/api/admin/limits/api_key/{id}/usage`) was always 0. Each
key now has counters of its own, rate limits and budgets alike, for the
gateway and the MCP gateway, and its usage shows what it used.
- **A refused request counts against nothing.** A request refused by a spent
budget, or by one rate limit after another had passed, was still counted
against the request limits. Budgets are now checked first and every rate
limit in one step, so a refused request leaves every counter as it was.
- **`Retry-After` says when to retry.** A `429` from the gateway's own limits
said `Retry-After: 30` whatever the limit. It now gives the seconds until the
window has room for another request, or until a spent budget's period ends
(the next midnight, Monday or 1st of the month, UTC). A spent budget also
sends `x-should-retry: false`, so the OpenAI and Anthropic SDKs don't retry it
by themselves. The body stays in the caller's API format.
- **Redis Cluster.** The rate-limit, route-health and quota scripts touched
keys of several hash slots, which a Redis Cluster refuses: on a cluster, rate
limits silently stopped applying (or refused every request with
`security.rate_limit_fail_closed`), circuit breakers never tripped, and cache
invalidation reached one node only. Every key a script touches now shares a
hash tag, pattern deletes scan every node, and the Helm chart's README
describes a `redis-cluster://` URL.
- **Redis over TLS.** The server was built without TLS for Redis: a
`rediss://` URL was used as plain TCP, so against a Redis that requires
TLS — ElastiCache with in-transit encryption, Upstash, Azure Cache for
Redis and Redis Cloud among them — the server did not start
(`Failed to connect to Redis: Protocol Error: Expected string.`).
`rediss://` and `rediss-cluster://` URLs now connect over TLS and check
the certificate against the system's CAs, as upstream HTTPS does. For a
Redis whose certificate a private CA signed, `REDIS_CA_CERT` names a PEM
file with that CA, which is then trusted alone; the Helm chart sets it
from a Secret given in `redis.caSecret`. The chart's README describes
both.
- **Several instances starting at once.** Server instances starting together
against one database — a Helm `replicaCount` above 1, a rolling upgrade, an
autoscaler adding pods — applied the schema side by side, and all but one
could exit with `Database migration failed: apply db/schema.sql: … deadlock
detected` (on an empty database: `duplicate key value violates unique
constraint "pg_extension_name_index"`). With ClickHouse, the rollups that an
instance fills from the logs when it finds them empty (`cost_rollup_hourly`,
`provider_health_5m`, `mcp_server_call_counts`) could be filled by each of
them, counting every request once per instance on the cost pages, the
dashboard and the MCP server list. Instances now set up Postgres and
ClickHouse one at a time, under Postgres advisory locks: the others wait,
logging `Another instance is setting up the database schema; waiting for it
to finish`, then find it done. An instance that dies holding a lock releases
it with its connection. The Helm chart's startup probe allows a pod 125 s
to start instead of 35 s, set in `startupProbe` (the chart's README says
when to raise it), and an attempt to reach a ClickHouse that doesn't answer
gives up after 5 s instead of the system's TCP timeout.
- **Captured bodies kept as long as configured.** With ClickHouse and
`audit.body_retention_days` above 30, every server start could clear the
captured request and response bodies older than 30 days
(`gateway_logs.request_body` / `response_body`, `mcp_logs.tool_arguments`
/ `tool_result`; the rows themselves stayed). The start-up table setup set
those columns' TTL to 30 days each time, and ClickHouse applies a TTL to
the data already stored as soon as it is set, before the server put the
configured TTL back a moment later. The setup now gives these columns a
TTL only when it creates them, so a restart leaves the configured one in
place. Bodies already cleared cannot be recovered. The log tables' own
TTLs (`data.retention_days_*`) were not affected.
- **Helm network policy and databases on other ports.** With
`networkPolicy.enabled`, the server could reach PostgreSQL only on `5432`,
Redis on `6379` and ClickHouse on `8123`, whatever their `externalUrl` said,
so a database on another port was blocked — Azure Cache for Redis over TLS
(`6380`), a managed Postgres or a ClickHouse on a port of its own: the
server could not start, or started without writing to ClickHouse. The
allowed ports now follow `postgres.externalUrl`, `redis.externalUrl` and
`clickhouse.externalUrl`: every port a URL names, and the client's default
for its scheme where it names none. `networkPolicy.extraEgress` adds egress
rules as written, for ports no URL names (Redis Cluster nodes announcing
other ports, an upstream or MCP server on a port other than `443`). The
chart's README describes both. Port `9000`, ClickHouse's native protocol,
is no longer allowed: the server reaches ClickHouse over plain HTTP only
(an `http://` URL; HTTPS is not supported). An S3 endpoint on `9000`
(RustFS, MinIO) configured outside the chart needs a rule in
`networkPolicy.extraEgress`.

## [3.1.0] — 2026-10-05

The thinkwatch-core crates move from v0.59.0 to v0.62.0. Two changes reach
Expand Down Expand Up @@ -1180,7 +1331,8 @@ unreleased builds should: stop the gateway, run `db/schema.sql`
against PostgreSQL, restart against this tag. The schema is
idempotent end-to-end, so the apply is safe to repeat.

[Unreleased]: https://github.com/ThinkWatchProject/ThinkWatch/compare/v3.1.0...HEAD
[Unreleased]: https://github.com/ThinkWatchProject/ThinkWatch/compare/v3.2.0...HEAD
[3.2.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.2.0
[3.1.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.1.0
[3.0.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.0.0
[2.2.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v2.2.0
Expand Down
20 changes: 14 additions & 6 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

11 changes: 9 additions & 2 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ members = [
]

[workspace.package]
version = "3.1.0"
version = "3.2.0"
edition = "2024"
# Pin the MSRV to the first stable rustc that ships edition 2024 (1.85,
# released 2025-02-20). Without this, contributors on older toolchains
Expand Down Expand Up @@ -90,7 +90,14 @@ rust_decimal = { version = "1", features = ["serde", "serde-with-str"] }
clickhouse = { version = "0.13", features = ["time", "chrono"] }

# Redis
fred = { version = "10", features = ["subscriber-client", "i-scripts"] }
# `enable-rustls` gives fred its rustls connector, for `rediss://` URLs.
# The connector itself is built in `think_watch_common::redis_config`
# with the same TLS stack the HTTP client uses (below): rustls on
# aws-lc-rs, the platform's roots through rustls-platform-verifier. No
# OpenSSL anywhere, so the static musl image needs nothing new.
fred = { version = "10", features = ["subscriber-client", "i-scripts", "enable-rustls"] }
rustls = { version = "0.23", default-features = false, features = ["aws_lc_rs", "std", "tls12"] }
rustls-platform-verifier = "0.6"

# Auth
# Crypto-critical crates are pinned with `=` so a silent minor-release
Expand Down
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,9 @@ The gateway (port `3000`) is the only part that clients need to reach. The conso
- A model's maximum output tokens, set on the Models page, caps `max_tokens` on every request to that model; it replaces the old output length guardrail.

**Limits and budgets**
- Request-count limits are checked before the request; token limits and budgets are counted after the response, so one request can cross a budget before the next is refused.
- Every limit and budget is checked before the request, against what earlier requests used; tokens are counted after the response, so one request can cross a token limit or a budget before the next is refused. A refused request counts against nothing, and its `429` says in `Retry-After` when the limit frees.
- A limit set on an API key applies on top of its owner's, on a counter of its own.
- Redis can be a single node or a Redis Cluster.
- If Redis is unavailable, limits fail open by default. Setting `security.rate_limit_fail_closed` refuses requests instead.
- Budget alerts fire once per period at 50%, 80%, 95% and 100%.

Expand Down
4 changes: 3 additions & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,9 @@ cd web && pnpm install && pnpm dev
- 模型的「最大输出 token」在模型页设置,限制发往该模型的每个请求的 `max_tokens`,取代原来的输出长度护栏。

**限流与预算**
- 请求数限制在请求发出前检查;Token 限制与预算在响应返回后计入,因此单个请求可能越过预算,此后的请求才会被拒绝。
- 所有限流与预算都在请求发出前按此前的用量检查;Token 在响应返回后计入,因此单个请求可能越过 Token 限制或预算,此后的请求才会被拒绝。被拒绝的请求不计入任何限制,其 `429` 响应以 `Retry-After` 说明限制何时解除。
- 设置在 API Key 上的限制叠加在其所属用户的限制之上,单独计数。
- Redis 可以是单节点,也可以是 Redis Cluster。
- Redis 不可用时,限流默认放行;设置 `security.rate_limit_fail_closed` 后改为拒绝请求。
- 预算提醒在每个周期内于 50%、80%、95% 和 100% 各触发一次。

Expand Down
Loading
Loading