Skip to content

Fix what the pre-release review of the guard unification found - #75

Merged
fylorn merged 3 commits into
devfrom
fix/guard-review
Oct 3, 2026
Merged

fylorn merged 3 commits into
devfrom
fix/guard-review

Conversation

@fylorn

@fylorn fylorn commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Fixes from the pre-release review of #72 against v2.2.0.

  • Convert once. The first start records the conversion in security.legacy_converted (time, version, what it converted). With the record present, old keys or the models.output_guardrails column showing up again — a 2.2 process started against the database writes its seeded defaults back — are removed with a warning, never converted over the policies in force. What the conversion leaves out (a rule that does not compile, an unknown built-in id, an unreadable output_guardrails, a renamed duplicate) is logged.
  • No request text in audit events. Content events name the rule, outcome, count and whether in a tool result — no excerpt, no revealed text. Redaction events are aggregated per rule per request with distinct values and occurrences; built-in rules keep core's masked form (up to 5), custom rules write no values. At most 20 events per guard per request, refusing/stripping hits first, each with rules_in_request. The refusal message to the caller is unchanged (masked).
  • Limits back in guard_policy::validate: content 500 custom rules / 500 characters per pattern; redaction 100 / 1,000; tool calls 100 / 1,000 and no custom name equal to a built-in id. 400 with the reason.
  • Write permission as in 2.2: PATCH of a security.* key takes settings:write and pii_redactor:write / content_filter:write. The console's write check (web/src/routes/gateway/security/index.tsx) asks for both; the trial endpoints are unchanged.
  • Output cap edge cases: applied per hop — a limit the request carries is lowered; a missing one is filled only when the cap is within fallback_max_output_tokens(upstream model) (32,000 Claude / 8,192 otherwise), otherwise the request goes out without one and the model's default applies. A converted cap is stored as it converts (ceil(N / 4)), not lowered to that figure: lowering it would also tighten requests that set their own limit, cutting answers 2.2 let through. The response cache key includes the cap. A model API request setting a non-empty output_guardrails is refused (400) instead of silently ignored. With core v0.59.0 (ThinkWatch-Core#279) the cap gets the hop's official flag, the one the conversion gets: a Chat request with no limit is given max_completion_tokens on OpenAI's own endpoint and max_tokens elsewhere, both Chat fields are lowered when both are set, and an Anthropic thinking budget is lowered with the cap. The model form's hint (maxOutputTokensHint, en/zh) says when a request without a limit gets the cap.
  • CHANGELOG "Read before upgrading": no rollback to 2.2 and backups, stopping 2.2 replicas (Helm commands), the conversion record, per-model cap check (what a request without a limit gets above the family figure, and how to hold it) and thinking tokens, stricter built-in PII rules, content events for empty lists and hidden_text: log, renamed metrics, PolicyBlocked error type, permissions, stripped bodies in the audit, output_guardrails refused. The Fixed entry no longer claims 2.2 stored content filter excerpts.

Core crates move to ThinkWatch-Core v0.59.0 (?tag=v0.59.0#ef10a8b; only the four core entries change in Cargo.lock).

Tests: unit tests for the field a Chat request without a limit gets on an official endpoint and elsewhere, both Chat limits held to the cap, the conversion record helpers, the cap stored as converted, the limits, per-rule aggregation and event ordering; integration tests for settings written back after the conversion, custom redaction events without values, the 20-event cap, the permission matrix, the limits, the output cap above the family figure and through an Anthropic conversion, and output_guardrails refused.

🤖 Generated with Claude Code

fylorn and others added 3 commits October 3, 2026 17:26
- Converting the old guard settings happens once. The first start
  records it in `security.legacy_converted`; old keys or the old column
  showing up after that (an older version started against the database
  writes its seeded defaults back) are removed with a warning, never
  converted over the policies in force. Whatever the conversion leaves
  out is logged.
- Audit events carry no text of the request. A content event names the
  rule and counts the matches. A redaction event is one per rule per
  request, with how many values and occurrences; only a built-in rule's
  lists a few in masked form, a custom rule's values are not written.
  At most 20 events per guard per request.
- Policy limits are back: 500 custom content rules of up to 500
  characters, 100 redaction and 100 tool-call rules of up to 1,000, and
  no custom tool-call rule named like a built-in one.
- Writing a guard policy takes `settings:write` and the guard's own
  permission, as 2.2 required across server and console; the console's
  write check asks for both.
- The output cap is applied per hop: a limit the request carries is
  lowered, and one it does not is filled in only within what the
  model's family is known to take. Converted caps are kept within it
  too. The cache key carries the cap. A model request still setting
  `output_guardrails` is refused instead of ignored.
- CHANGELOG: what to know before upgrading (no way back to 2.2, backups,
  stopping replicas under Helm, stricter built-in PII rules, renamed
  metrics, error type, permissions), and the Fixed entry corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The upgrade lowered a converted cap to the 8,192 tokens (32,000 for
Claude) the gateway fills in for a request that sets no limit. That
lowered the limit of every request that set one above it too: a
100,000-byte cap on a non-Claude model, 25,000 tokens, started cutting
answers at 8,192, a tightening on upgrade alone.

The cap is now stored as it converts. The runtime rule stays: a
request's own limit is lowered to the cap, and a request with none gets
the cap only when it is within that family figure. The changelog says
what that means for requests that set no limit, and how to hold them to
a cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
thinkwatch-core v0.59.0 (ThinkWatch-Core#279) fixes three edge cases of
cap_max_output_tokens, which now takes whether the hop goes to the
vendor's own endpoint:

- a Chat request that sets both max_tokens and max_completion_tokens has
  both lowered, not only the one the decoder reads first;
- one that sets neither is given max_completion_tokens on OpenAI's own
  endpoint, whose reasoning models refuse max_tokens, and max_tokens
  elsewhere;
- an Anthropic request with extended thinking has its budget lowered
  below the cap too, or thinking dropped when the cap is 1,024 tokens or
  less.

It also makes outbound redaction linear in the number of secrets in a
request. The gateway passes the same official flag the conversion gets.
The model form's hint and the changelog now say when a request that
sets no limit gets the cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit 2baa123 into dev Oct 3, 2026
6 checks passed
@fylorn
fylorn deleted the fix/guard-review branch October 3, 2026 10:39
@fylorn fylorn mentioned this pull request Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant