Skip to content

Unify the three guards on one shared model (protocol 33, SCHEMA 24) - #268

Merged
fylorn merged 6 commits into
mainfrom
feat/guard-unify
Oct 2, 2026
Merged

fylorn merged 6 commits into
mainfrom
feat/guard-unify

Conversation

@fylorn

@fylorn fylorn commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Unifies the three guards (outbound redaction, tool-call inspection, content filter) on one model defined in tw-guard, so the desktop app and the enterprise edition share the policy shape, built-in catalogs, validation, rule views and the "test" logic. The desktop app writes the policy to config.yaml; the enterprise edition stores the same structure in system settings.

Commits:

  • Part 1 (16c451e): shared layer in tw-guard / tw-dialect (policy, view, trial, content engine with strip and code point rules, redaction flow, caller-text walker).
  • Add two built-in tool-call rules for credential and file exfiltration #266 (93ff769): two code-backed tool-call rules (secret-to-unknown-host, upload-file-to-host).
  • Part 2 (d86320b): desktop config, gateway flow and control API on the shared model.
  • 73d9328: why legacy completions (/v1/completions, /v1/complete) are not content-screened.
  • Say which rule a cut tool call matched, not who produced it #269 (414759e): the tool-call cut messages name the rule, not who produced the call (a plugin can produce one too).
  • a20ca47: README (three protections).

Breaking changes

Config

  • security.hidden_text and security.output_limit are removed. A config.yaml that still has either key does not load (the app starts in safe mode) until the key is deleted. No migration, by design.
  • Hidden characters are built-in content rules: unicode-tags, bidi-controls (on), zero-width, private-use (off).
  • Content rule actions are block | strip | record (warn / log removed). Custom content rules can match by codepoints; custom redaction rules accept label (placeholder <<TW_{LABEL}_n>>).

Gateway behaviour

  • Content screening runs before the request starts. Stripped text is removed from user messages and tool results, and the stripped body (plus the IR decoded again from it) is what gets redacted, recorded and sent on every hop. Refused requests still leave a failed row (denied).
  • Compaction requests (/v1/responses/compact, /backend-api/codex/responses/compact) are now screened like generation requests. Token-count requests are deliberately not screened.
  • WebSocket frames: response.create frames are screened and stripped; other frames go through code point rules only.
  • The output limit is removed everywhere.
  • Fix: tool-call excerpts are masked with the stored-body redaction before tool_call_flagged is emitted. A secret restored from a placeholder used to reach the event, the security log and system notifications in the clear.

Control API: CONTROL_API_VERSION = 33

  • Guard is redact | inspect_tools | content; SecurityDetail / SecurityView have those three fields.
  • PUT /security/{guard}/limit (SetSecurityLimit) is removed.
  • Types come from tw-guard: SecurityRuleView gains label; Matcher gains email, cn-mobile-phone, builtin; RuleAction gains strip; ContentMatch gains codepoints; CustomRuleSave gains label; SecurityTestRequest gains label and action; SecurityTestResult gains output and refused.
  • Events: hidden_text_found and output_limited are removed; content_matched now has match, action, outcome (recorded | stripped | blocked), count and revealed (no blocked flag).
  • SecurityOutcome and SecurityOutcomeCounts gain stripped; SecurityEventView gains match and revealed; /summary security counts are secrets, secrets_replaced, tool_calls, tool_calls_cut, content, content_blocked, content_stripped.

Request store: SCHEMA = 24

The security log gains match and revealed. Request history is reset on upgrade.

Message codes

  • Added: config.rule_codepoints_bad (name, detail), config.rule_label_bad (name, label), gw.content.refused_invisible_message / gw.content.refused_invisible_tool_result (rule, name, count), security.bad_codepoints (detail), security.bad_label (label), security.content_action_unknown (action), security.pattern_empty, security.unknown_guard (guard).
  • Renamed (Say which rule a cut tool call matched, not who produced it #269, same parameters unless noted): gw.toolcall.cut → gw.toolcall.response_cut, gw.toolcall.blocked → gw.toolcall.response_withheld, gw.ws.toolcall_cut → gw.toolcall.connection_cut (detail → why).
  • Removed: config.output_limit_range, gw.hidden_text.refused_message, gw.hidden_text.refused_tool_result, gw.output_limit.cut, gw.output_limit.withheld, security.guard_unknown, security.limit_range, security.no_custom_rules, security.no_limit, security.nothing_to_test, security.unknown_content_action.

Rust API (layer one, used by the enterprise edition)

  • tw_guard::output is deleted. In tw_guard::hidden only the config-file scan remains. content::Action::Warn / Log and content::scan_request / worst are gone; use tw_guard::content::screen and tw_guard::policy.
  • The enterprise CI job builds enterprise feat/guard-unify (Unify the request guards with thinkwatch-core's rule model ThinkWatch#72), which moves to the shared model, and passes.

Checks

cargo fmt --check, cargo clippy --workspace --all-targets -D warnings (plus tw-api / tw-guard with ts), and cargo test --workspace: 2181 passed.

Other sides

🤖 Generated with Claude Code

fylorn and others added 6 commits October 3, 2026 00:45
…engines

Add the layer both gateways will share for outbound redaction, tool-call
inspection and content filtering. Nothing is removed yet: the desktop keeps
its current configuration, rule catalog and behaviour; the next part switches
it over and deletes the old interfaces.

tw-guard
- policy: Mode, Guard, ToolAction, ContentAction (block | strip | record),
  ContentMatch (contains | regex | codepoints), the three policies and their
  custom rules (redaction rules take an optional placeholder label), one
  check() with stable error codes, and rules()/one_builtin() compilers.
  Same shape in config.yaml and in JSON.
- content: code point matching (U+200B, U+E0000–U+E007F; up to 32 items),
  the Strip and Record actions, an "invisible" built-in group first in the
  catalog (Unicode tags and bidi controls on, zero-width and private-use
  off, all strip), Hit::count and Hit::revealed, and screen()/screen_text():
  find on the request's own JSON, delete every match of strip rules from the
  caller's text, look again after deleting (a keyword split by zero-width
  characters is caught once they are gone), refuse on block rules.
- redact::flow: hits/find/look/replace/ledger_for, moved from the desktop
  gateway, plus look_from (numbering on from a seed ledger). Hits inside
  base64 payloads (data URIs, image data, signatures) no longer count. The
  desktop's guard functions keep their names and signatures and call it.
- redact::rules: email and Chinese mainland mobile number built-ins, off by
  default, with their own placeholder labels and masking.
- view and trial: the rule view and the "try it" request/result both
  products return, with TypeScript export behind a `ts` feature. The view
  is lossless, so a policy can be rebuilt from it.

tw-dialect
- caller: where the caller's text lives in each client format (Anthropic,
  Chat, Responses, Gemini, Bedrock), as positions that can be read and
  rewritten; a comparison test pins it to what the decoders read.
- params: read and write the maximum output tokens, the model and the
  thinking switch per format. The desktop's routing `set` now uses it.

tw-api re-exports the shared types under `tw_api::guard`; the endpoints
still use the old ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#266)

A malicious upstream can write a tool call that, once the gateway restores
redaction placeholders to their real values, sends a credential somewhere it
should not go -- e.g. a curl to an unknown host with the key in the query.
Redaction and tool-call inspection are separate jobs: redaction keeps secrets
off the wire, and the restored, final tool call the client will execute is
judged by the tool-call guard. Close the gap with two rules in the "dangerous"
group.

- secret-to-unknown-host (high, cut on enforce): the tool call makes an
  http(s) request and its arguments carry a value the outbound-redaction
  detector recognizes as a credential (API key or private key), while the
  destination is neither local (loopback, localhost, *.localhost) nor the
  credential's own provider. The provider map is small, explicit and keyed by
  the redaction rule ids.
- upload-file-to-host (medium, record only): the tool call uploads a local
  file's contents to a non-local host (curl -T / --data @file / -F field=@file
  / --upload-file / --post-file). Common in development, so it only records.

Both are code-backed: the decision spans the arguments (URL plus credential
plus upload marker) and cannot be one regex. RuleSpec gains an optional
`check` field; code-backed rules get a never-matching `re` and stay out of
scan_rules(), so the client-config scanner and anything still reading
`rule.re` directly never match them. Rule::find() handles both kinds, and the
wall, the trial and the desktop test endpoint go through it. The rule view
gains a Matcher::Builtin { check } variant so the security page can list them.

The rules see the full, restored arguments across streaming, non-streaming and
WebSocket answers via the existing wall accumulation; work per call is bounded.
Config docs regenerated.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The desktop side now runs on the shared policy model from part 1: one
definition of the three guards (outbound redaction, tool-call inspection,
content filter) in tw-guard, written to config.yaml here and to system
settings in the enterprise edition.

Config (breaking):
- `security` is `tw_guard::policy::Security`. `security.hidden_text` and
  `security.output_limit` are gone; a config that still has them does not
  load. Hidden characters are content rules now (built-in `unicode-tags`,
  `bidi-controls`, plus `zero-width` and `private-use`, off by default).
- Content rules act by `block`, `strip` or `record`; `warn`/`log` are gone.
  Custom content rules can match by `codepoints`; custom redaction rules
  take a placeholder `label`.
- Policy errors map to config message codes; the interim config module,
  `tw_guard::output` and the request-only parts of `tw_guard::hidden` are
  deleted. Config manual regenerated.

Gateway:
- The content filter screens a request before it starts. When rules strip
  text, the stripped body (and the IR decoded again from it) is what gets
  redacted, recorded and sent on every hop. A refused request still leaves
  a failed row in traffic. Findings are reported on the request id.
- Compaction requests (`/v1/responses/compact`, Codex
  `/backend-api/codex/responses/compact`) are screened like generation
  requests: they carry the whole conversation and run a model. Token counts
  are not screened: no model runs, and screening them would record the same
  finding twice and could refuse the client's count.
- WebSocket frames: decodable `response.create` frames are screened and
  stripped like HTTP requests; other frames go through code point rules.
- The output limit is removed from every path.
- Fix: the excerpt of a flagged tool call is masked with the same redaction
  as stored bodies before it goes into `tool_call_flagged`. A secret
  restored from a placeholder used to reach the event, the security log
  and the system notification in the clear.

Control API (protocol 33, request store SCHEMA 24, history is reset):
- `Guard` is `redact | inspect_tools | content`; `SecurityDetail` and
  `SecurityView` have those three; `SetSecurityLimit` is removed.
- Rule views, the test endpoint and their types come from tw-guard
  (`tw_guard::view`, `tw_guard::trial`); tests return `output` and
  `refused`. `SecurityRuleView` has `label`; `RuleAction` has `strip`;
  `ContentMatch` has `codepoints`; `CustomRuleSave` has `label`.
- `content_matched` carries `match`, `action`, `outcome`
  (`recorded | stripped | blocked`) and `revealed`; `hidden_text_found` and
  `output_limited` are gone. `SecurityOutcome` gains `stripped`; the log
  entry gains `match` and `revealed`; summary counts are
  `secrets, secrets_replaced, tool_calls, tool_calls_cut, content,
  content_blocked, content_stripped`.

Message codes added: config.rule_codepoints_bad, config.rule_label_bad,
gw.content.refused_invisible_message, gw.content.refused_invisible_tool_result,
security.bad_codepoints, security.bad_label, security.content_action_unknown,
security.pattern_empty, security.unknown_guard. Removed:
config.output_limit_range, gw.hidden_text.refused_message,
gw.hidden_text.refused_tool_result, gw.output_limit.cut,
gw.output_limit.withheld, security.guard_unknown, security.limit_range,
security.no_custom_rules, security.no_limit, security.nothing_to_test,
security.unknown_content_action.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
/v1/completions and /v1/complete carry one prompt string, with no way to
tell the caller's text from a tool result, and today they mostly serve
editor code completion, where there are no tool results to carry an
injection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tool-call guard judges the answer as the client would receive it,
and a script plugin with the tool-call permission can produce or modify
tool calls. "The call returned by upstream X" is no longer always true,
so the sentences now say the answer contained the call.

New message codes (the old ones are removed):
- gw.toolcall.cut -> gw.toolcall.response_cut
- gw.toolcall.blocked -> gw.toolcall.response_withheld
- gw.ws.toolcall_cut -> gw.toolcall.connection_cut

All three take the same arguments: upstream, tool, rule, name, why (the
WebSocket code used to call the reason `detail`). The upstream stays an
argument, as it stays in tool_call_flagged and in the log fields.

A WebSocket connection cut for a tool call now tells the client the same
coded sentence its ending records, as a frame the content filter refuses
already does, instead of a separate hard-coded line. The log lines no
longer name the upstream as the source either.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ructions

The output limit is gone and hidden characters are content filter rules,
so the highlights and the crate table describe three protections:
outbound redaction, tool-call inspection and the content filter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn marked this pull request as ready for review October 2, 2026 20:14
@fylorn
fylorn merged commit 72954bf into main Oct 2, 2026
4 checks passed
@fylorn
fylorn deleted the feat/guard-unify branch October 2, 2026 20:15
fylorn added a commit that referenced this pull request Oct 3, 2026
Release for #251, #252, #263, #265, #268 and #272, with the pre-release
fixes in #273.

- Bump the workspace version to 0.58.0
- release-notes/0.58.0.md

The control-plane protocol (34) and the request store schema (25) are
already at their release values on main and do not change here.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant