Skip to content

Merge main into feat/plugins: plugins on the unified guards - #271

Merged
fylorn merged 5 commits into
feat/pluginsfrom
plugins-sync-main
Oct 2, 2026
Merged

fylorn merged 5 commits into
feat/pluginsfrom
plugins-sync-main

Conversation

@fylorn

@fylorn fylorn commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Merges current main (#263, #265, #268) into feat/plugins, so the plugins land on top of the unified guards. Merged with a merge commit; feat/plugins gets fast-forwarded to this branch so main's history stays linked.

Conflicts and decisions

  • Versions: CONTROL_API_VERSION 34 and request store SCHEMA 25. Both branches had used 33 and 24.
    • The version notes keep 33 for the guard unification and describe the plugins (UpdatePluginConfirmed, request kinds, PluginRunView.attempt, …) under 34.
    • New tests: the version must be the newest one its notes describe, and either kind of version-24 database (without plugin_runs, or without the new security_events columns) is rebuilt. Both tests fail on 33/24.
  • guard.rs: main's flow-based functions. The plugin-only screen_more is replaced by rescreen/added on the new Screening model.
  • Pipeline:
    • Main's step-4 screening (after routing, before the start event) strips the client's original.
    • The per-attempt request hook runs on that body.
    • plug.rs re-screens a changed request with guard::rescreen.
  • relay.rs: the output meter is gone. Reply hooks stay after conversion and before wall_cut. The whole-answer-as-stream wall path from feat/plugins now uses gw.toolcall.response_withheld and masks excerpts.
  • ws.rs:
    • Request side: main's screen_frame (with strip), then plugin_request on the screened frame, then redaction.
    • Reply side: hooks, then the wall with gw.toolcall.connection_cut.
    • The per-answer response.failed (response, dropping, failed_frame) is kept for plugin failures only.
  • state.rs, tw-control/security.rs, db.rs, msg-codes.txt, m5_toolwall.rs: both sides kept. Message codes were regenerated, not hand-merged.

Re-screening plugin output

  • Only what a plugin added is reported, and only that can refuse.
  • What it added that strip rules match is stripped before conversion, redaction and sending.
  • Token counts are still unscreened, but plugins still run on them.
  • Embeddings and legacy completions are not screened up front. The inputs a plugin changed are screened as caller text and stripped in place; untouched inputs are left alone.
  • Hits the plugin's input already had are not reported again. Embeddings and completions can carry block-rule hits because they are not screened up front, so those hits cannot refuse a request the plugin only partly changed.

Wasmtime 49.0.2

The Audit check failed on three advisories published against Wasmtime 49.0.1, which runs the plugin sandbox (RUSTSEC-2026-0325, -0326, -0327). One extra commit pins both the runtime and the build-time compiler to 49.0.2, which fixes all three; Cranelift 0.136.2 and wasmparser/wasm-encoder 0.258.3 come with it.

Tests

  • Removed: two plugin tests for the deleted output limit.
  • Changed: the hidden-character plugin test now expects the built-in unicode-tags rule to strip what a plugin hides, and to refuse when its action is block.
  • Added:
    • unit tests for rescreen
    • an end-to-end test that a zero-width character a plugin writes into one embeddings input is stripped there and nowhere else
    • the two version tests

Local: cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo test --workspace (2588 passed), and tw-api with ts plus tsc.

The enterprise job checks against ThinkWatch dev, which still fails on main since #268 until ThinkWatch#72 lands. This PR leaves the layer-one crates byte-identical to main.

🤖 Generated with Claude Code

fylorn and others added 5 commits October 2, 2026 23:08
…arries the status (#263)

When an upstream answered 4xx (or 3xx) and the gateway handed the answer
on to the client, the ending was RequestFinished with that status and an
empty error. Everything that counts failures reads `error`, so the
request looked successful everywhere: the traffic list, the overview's
failed count, the session's errors, the upstream check-up, the history
search's failed filter. Session detail (TurnView) treated the turn as
done while the transcript, which reads the status, gave it no answer:
the desktop app's conversation showed only the user's message, with no
answer and no reason.

The gateway now reports these as RequestFailed (source `upstream`):

- relay marks a non-2xx answer on the Ending (`Ending::refused`). The
  Ending keeps the head of the error body whether or not bodies are
  stored, and on `finished` reports the failure with what the upstream
  said: gw.upstream.status_message {upstream, status, message}
  ("Upstream `x` answered 400: prompt is too long..."), read in the
  upstream's format with tw_dialect::convert::error_message (now shared
  with Session::error). Plain text is used as it is; an empty body or an
  HTML page falls back to gw.upstream.status. The words are masked with
  the request's redaction, like the stored response body, and capped at
  500 characters.
- A Bedrock credential refusal already replaces AWS's words with ours;
  that message is now the recorded reason. The `[ThinkWatch]` prefix
  moves out of those four messages into the body sent to the client, as
  with the gateway's own errors.
- A client that leaves while the error is being passed on is still a
  cancellation. WebSocket upgrades (101) are unchanged.

The row's status still comes from RequestHeaders, so one rule holds in
every place that counts: a request failed if and only if `error` is set.
No SQL changes, and cancellations are counted as before.

TurnView gains `status` (the upstream's status code, null when the
request never got one), so a failed turn can say what the upstream
answered. CONTROL_API_VERSION stays at 32 (unreleased); its note
mentions both changes.

Tests: the Ending (the upstream's words, nothing readable, masking, our
words for Bedrock, a cancellation during an error answer); a passed-on
400 reaches the client unchanged and ends once as a failure; the Bedrock
refusal's recorded reason; rows, turns, session and summary in the
recorder; and an end-to-end control-plane test through a real gateway
and store with a 200, a passed-on 400 and a cancelled turn in one
session, read back from /sessions/{id}, /history, /summary, /sessions
and the transcript.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The advisories job on the plugins branch flagged rustls 0.23.44;
main still had it.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…268)

* Guard unify, part 1: the shared policy model, rule views, trials and engines

Add the layer both gateways will share for outbound redaction, tool-call
inspection and content filtering. Nothing is removed yet: the desktop keeps
its current configuration, rule catalog and behaviour; the next part switches
it over and deletes the old interfaces.

tw-guard
- policy: Mode, Guard, ToolAction, ContentAction (block | strip | record),
  ContentMatch (contains | regex | codepoints), the three policies and their
  custom rules (redaction rules take an optional placeholder label), one
  check() with stable error codes, and rules()/one_builtin() compilers.
  Same shape in config.yaml and in JSON.
- content: code point matching (U+200B, U+E0000–U+E007F; up to 32 items),
  the Strip and Record actions, an "invisible" built-in group first in the
  catalog (Unicode tags and bidi controls on, zero-width and private-use
  off, all strip), Hit::count and Hit::revealed, and screen()/screen_text():
  find on the request's own JSON, delete every match of strip rules from the
  caller's text, look again after deleting (a keyword split by zero-width
  characters is caught once they are gone), refuse on block rules.
- redact::flow: hits/find/look/replace/ledger_for, moved from the desktop
  gateway, plus look_from (numbering on from a seed ledger). Hits inside
  base64 payloads (data URIs, image data, signatures) no longer count. The
  desktop's guard functions keep their names and signatures and call it.
- redact::rules: email and Chinese mainland mobile number built-ins, off by
  default, with their own placeholder labels and masking.
- view and trial: the rule view and the "try it" request/result both
  products return, with TypeScript export behind a `ts` feature. The view
  is lossless, so a policy can be rebuilt from it.

tw-dialect
- caller: where the caller's text lives in each client format (Anthropic,
  Chat, Responses, Gemini, Bedrock), as positions that can be read and
  rewritten; a comparison test pins it to what the decoders read.
- params: read and write the maximum output tokens, the model and the
  thinking switch per format. The desktop's routing `set` now uses it.

tw-api re-exports the shared types under `tw_api::guard`; the endpoints
still use the old ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add two built-in tool-call rules for credential and file exfiltration (#266)

A malicious upstream can write a tool call that, once the gateway restores
redaction placeholders to their real values, sends a credential somewhere it
should not go -- e.g. a curl to an unknown host with the key in the query.
Redaction and tool-call inspection are separate jobs: redaction keeps secrets
off the wire, and the restored, final tool call the client will execute is
judged by the tool-call guard. Close the gap with two rules in the "dangerous"
group.

- secret-to-unknown-host (high, cut on enforce): the tool call makes an
  http(s) request and its arguments carry a value the outbound-redaction
  detector recognizes as a credential (API key or private key), while the
  destination is neither local (loopback, localhost, *.localhost) nor the
  credential's own provider. The provider map is small, explicit and keyed by
  the redaction rule ids.
- upload-file-to-host (medium, record only): the tool call uploads a local
  file's contents to a non-local host (curl -T / --data @file / -F field=@file
  / --upload-file / --post-file). Common in development, so it only records.

Both are code-backed: the decision spans the arguments (URL plus credential
plus upload marker) and cannot be one regex. RuleSpec gains an optional
`check` field; code-backed rules get a never-matching `re` and stay out of
scan_rules(), so the client-config scanner and anything still reading
`rule.re` directly never match them. Rule::find() handles both kinds, and the
wall, the trial and the desktop test endpoint go through it. The rule view
gains a Matcher::Builtin { check } variant so the security page can list them.

The rules see the full, restored arguments across streaming, non-streaming and
WebSocket answers via the existing wall accumulation; work per call is bounded.
Config docs regenerated.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Guard unify, part 2: config, gateway and control API on the shared model

The desktop side now runs on the shared policy model from part 1: one
definition of the three guards (outbound redaction, tool-call inspection,
content filter) in tw-guard, written to config.yaml here and to system
settings in the enterprise edition.

Config (breaking):
- `security` is `tw_guard::policy::Security`. `security.hidden_text` and
  `security.output_limit` are gone; a config that still has them does not
  load. Hidden characters are content rules now (built-in `unicode-tags`,
  `bidi-controls`, plus `zero-width` and `private-use`, off by default).
- Content rules act by `block`, `strip` or `record`; `warn`/`log` are gone.
  Custom content rules can match by `codepoints`; custom redaction rules
  take a placeholder `label`.
- Policy errors map to config message codes; the interim config module,
  `tw_guard::output` and the request-only parts of `tw_guard::hidden` are
  deleted. Config manual regenerated.

Gateway:
- The content filter screens a request before it starts. When rules strip
  text, the stripped body (and the IR decoded again from it) is what gets
  redacted, recorded and sent on every hop. A refused request still leaves
  a failed row in traffic. Findings are reported on the request id.
- Compaction requests (`/v1/responses/compact`, Codex
  `/backend-api/codex/responses/compact`) are screened like generation
  requests: they carry the whole conversation and run a model. Token counts
  are not screened: no model runs, and screening them would record the same
  finding twice and could refuse the client's count.
- WebSocket frames: decodable `response.create` frames are screened and
  stripped like HTTP requests; other frames go through code point rules.
- The output limit is removed from every path.
- Fix: the excerpt of a flagged tool call is masked with the same redaction
  as stored bodies before it goes into `tool_call_flagged`. A secret
  restored from a placeholder used to reach the event, the security log
  and the system notification in the clear.

Control API (protocol 33, request store SCHEMA 24, history is reset):
- `Guard` is `redact | inspect_tools | content`; `SecurityDetail` and
  `SecurityView` have those three; `SetSecurityLimit` is removed.
- Rule views, the test endpoint and their types come from tw-guard
  (`tw_guard::view`, `tw_guard::trial`); tests return `output` and
  `refused`. `SecurityRuleView` has `label`; `RuleAction` has `strip`;
  `ContentMatch` has `codepoints`; `CustomRuleSave` has `label`.
- `content_matched` carries `match`, `action`, `outcome`
  (`recorded | stripped | blocked`) and `revealed`; `hidden_text_found` and
  `output_limited` are gone. `SecurityOutcome` gains `stripped`; the log
  entry gains `match` and `revealed`; summary counts are
  `secrets, secrets_replaced, tool_calls, tool_calls_cut, content,
  content_blocked, content_stripped`.

Message codes added: config.rule_codepoints_bad, config.rule_label_bad,
gw.content.refused_invisible_message, gw.content.refused_invisible_tool_result,
security.bad_codepoints, security.bad_label, security.content_action_unknown,
security.pattern_empty, security.unknown_guard. Removed:
config.output_limit_range, gw.hidden_text.refused_message,
gw.hidden_text.refused_tool_result, gw.output_limit.cut,
gw.output_limit.withheld, security.guard_unknown, security.limit_range,
security.no_custom_rules, security.no_limit, security.nothing_to_test,
security.unknown_content_action.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Say why legacy completions are not content-screened

/v1/completions and /v1/complete carry one prompt string, with no way to
tell the caller's text from a tool result, and today they mostly serve
editor code completion, where there are no tool results to carry an
injection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Say which rule a cut tool call matched, not who produced it (#269)

The tool-call guard judges the answer as the client would receive it,
and a script plugin with the tool-call permission can produce or modify
tool calls. "The call returned by upstream X" is no longer always true,
so the sentences now say the answer contained the call.

New message codes (the old ones are removed):
- gw.toolcall.cut -> gw.toolcall.response_cut
- gw.toolcall.blocked -> gw.toolcall.response_withheld
- gw.ws.toolcall_cut -> gw.toolcall.connection_cut

All three take the same arguments: upstream, tool, rule, name, why (the
WebSocket code used to call the reason `detail`). The upstream stays an
argument, as it stays in tool_call_flagged and in the log fields.

A WebSocket connection cut for a tool call now tells the client the same
coded sentence its ending records, as a frame the content filter refuses
already does, instead of a separate hard-coded line. The log lines no
longer name the upstream as the source either.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* README: three protections, and the content filter removes hidden instructions

The output limit is gone and hidden characters are content filter rules,
so the highlights and the crate table describe three protections:
outbound redaction, tool-call inspection and the content filter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Brings #263 (an upstream error answer is a failed request), #265 (rustls
0.23.45) and #268 (the three guards on one shared model) into the plugin
branch. The output-length guard is gone and hidden characters are content
rules, so plugins now work with the content filter's block, strip and
record actions.

Request side, on every path (HTTP, compaction, WebSocket):
- The content filter screens the client's original once, where main does
  (after routing, before the start event), and strips it under enforce.
- Each upstream attempt runs its plugins on that body, stripped or not.
- A request a plugin changed is screened again (guard::rescreen): only what
  the plugin added is reported and only that can refuse; whatever it added
  that a strip rule matches is stripped before conversion, redaction and
  sending.
- Token counts are still not screened, but plugins still run on them.
- Embeddings and legacy completions are not screened up front; the inputs
  a plugin changed are screened as caller text and stripped in place.

Reply side: hooks stay between format conversion and the tool-call guard.
A cut uses the source-neutral codes (gw.toolcall.response_cut,
response_withheld, connection_cut) whoever produced the call, and flagged
excerpts are masked. On the WebSocket, response.failed for one answer is
kept for plugin failures only.

Versions: CONTROL_API_VERSION 34 (33 is the guard unification, 34 adds
plugins) and request store SCHEMA 25 (both sides had used 24 for different
tables). New tests pin both: the API version must be the newest one its
notes describe, and either kind of version-24 database is rebuilt.

Plugin tests for the deleted output limit are removed. The hidden-character
test now expects the built-in unicode-tags rule to strip what a plugin
hides, and to refuse when that rule's action is block. Message codes are
regenerated; the config manual and the default plugin manifests were
already current.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three advisories published against Wasmtime 49.0.1, which runs the plugin
sandbox, fail the Audit check. 49.0.2 fixes all three. The runtime and the
build-time compiler stay pinned to the same exact version, since a
precompiled module only loads in the Wasmtime that compiled it. Cranelift
moves to 0.136.2 and wasmparser/wasm-encoder to 0.258.3 with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit e8a19ca into feat/plugins Oct 2, 2026
4 of 5 checks passed
@fylorn
fylorn deleted the plugins-sync-main branch October 2, 2026 22:12
@fylorn fylorn mentioned this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant