Add two built-in tool-call rules for credential and file exfiltration - #266
Merged
Merged
Conversation
A malicious upstream can write a tool call that, once the gateway restores redaction placeholders to their real values, sends a credential somewhere it should not go -- e.g. a curl to an unknown host with the key in the query. Redaction and tool-call inspection are separate jobs: redaction keeps secrets off the wire, and the restored, final tool call the client will execute is judged by the tool-call guard. Close the gap with two rules in the "dangerous" group. - secret-to-unknown-host (high, cut on enforce): the tool call makes an http(s) request and its arguments carry a value the outbound-redaction detector recognizes as a credential (API key or private key), while the destination is neither local (loopback, localhost, *.localhost) nor the credential's own provider. The provider map is small, explicit and keyed by the redaction rule ids. - upload-file-to-host (medium, record only): the tool call uploads a local file's contents to a non-local host (curl -T / --data @file / -F field=@file / --upload-file / --post-file). Common in development, so it only records. Both are code-backed: the decision spans the arguments (URL plus credential plus upload marker) and cannot be one regex. RuleSpec gains an optional `check` field; code-backed rules get a never-matching `re` and stay out of scan_rules(), so the client-config scanner and anything still reading `rule.re` directly never match them. Rule::find() handles both kinds, and the wall, the trial and the desktop test endpoint go through it. The rule view gains a Matcher::Builtin { check } variant so the security page can list them. The rules see the full, restored arguments across streaming, non-streaming and WebSocket answers via the existing wall accumulation; work per call is bounded. Config docs regenerated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Merging into the integration branch with the |
fylorn
added a commit
that referenced
this pull request
Oct 2, 2026
…268) * Guard unify, part 1: the shared policy model, rule views, trials and engines Add the layer both gateways will share for outbound redaction, tool-call inspection and content filtering. Nothing is removed yet: the desktop keeps its current configuration, rule catalog and behaviour; the next part switches it over and deletes the old interfaces. tw-guard - policy: Mode, Guard, ToolAction, ContentAction (block | strip | record), ContentMatch (contains | regex | codepoints), the three policies and their custom rules (redaction rules take an optional placeholder label), one check() with stable error codes, and rules()/one_builtin() compilers. Same shape in config.yaml and in JSON. - content: code point matching (U+200B, U+E0000–U+E007F; up to 32 items), the Strip and Record actions, an "invisible" built-in group first in the catalog (Unicode tags and bidi controls on, zero-width and private-use off, all strip), Hit::count and Hit::revealed, and screen()/screen_text(): find on the request's own JSON, delete every match of strip rules from the caller's text, look again after deleting (a keyword split by zero-width characters is caught once they are gone), refuse on block rules. - redact::flow: hits/find/look/replace/ledger_for, moved from the desktop gateway, plus look_from (numbering on from a seed ledger). Hits inside base64 payloads (data URIs, image data, signatures) no longer count. The desktop's guard functions keep their names and signatures and call it. - redact::rules: email and Chinese mainland mobile number built-ins, off by default, with their own placeholder labels and masking. - view and trial: the rule view and the "try it" request/result both products return, with TypeScript export behind a `ts` feature. The view is lossless, so a policy can be rebuilt from it. tw-dialect - caller: where the caller's text lives in each client format (Anthropic, Chat, Responses, Gemini, Bedrock), as positions that can be read and rewritten; a comparison test pins it to what the decoders read. - params: read and write the maximum output tokens, the model and the thinking switch per format. The desktop's routing `set` now uses it. tw-api re-exports the shared types under `tw_api::guard`; the endpoints still use the old ones. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add two built-in tool-call rules for credential and file exfiltration (#266) A malicious upstream can write a tool call that, once the gateway restores redaction placeholders to their real values, sends a credential somewhere it should not go -- e.g. a curl to an unknown host with the key in the query. Redaction and tool-call inspection are separate jobs: redaction keeps secrets off the wire, and the restored, final tool call the client will execute is judged by the tool-call guard. Close the gap with two rules in the "dangerous" group. - secret-to-unknown-host (high, cut on enforce): the tool call makes an http(s) request and its arguments carry a value the outbound-redaction detector recognizes as a credential (API key or private key), while the destination is neither local (loopback, localhost, *.localhost) nor the credential's own provider. The provider map is small, explicit and keyed by the redaction rule ids. - upload-file-to-host (medium, record only): the tool call uploads a local file's contents to a non-local host (curl -T / --data @file / -F field=@file / --upload-file / --post-file). Common in development, so it only records. Both are code-backed: the decision spans the arguments (URL plus credential plus upload marker) and cannot be one regex. RuleSpec gains an optional `check` field; code-backed rules get a never-matching `re` and stay out of scan_rules(), so the client-config scanner and anything still reading `rule.re` directly never match them. Rule::find() handles both kinds, and the wall, the trial and the desktop test endpoint go through it. The rule view gains a Matcher::Builtin { check } variant so the security page can list them. The rules see the full, restored arguments across streaming, non-streaming and WebSocket answers via the existing wall accumulation; work per call is bounded. Config docs regenerated. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Guard unify, part 2: config, gateway and control API on the shared model The desktop side now runs on the shared policy model from part 1: one definition of the three guards (outbound redaction, tool-call inspection, content filter) in tw-guard, written to config.yaml here and to system settings in the enterprise edition. Config (breaking): - `security` is `tw_guard::policy::Security`. `security.hidden_text` and `security.output_limit` are gone; a config that still has them does not load. Hidden characters are content rules now (built-in `unicode-tags`, `bidi-controls`, plus `zero-width` and `private-use`, off by default). - Content rules act by `block`, `strip` or `record`; `warn`/`log` are gone. Custom content rules can match by `codepoints`; custom redaction rules take a placeholder `label`. - Policy errors map to config message codes; the interim config module, `tw_guard::output` and the request-only parts of `tw_guard::hidden` are deleted. Config manual regenerated. Gateway: - The content filter screens a request before it starts. When rules strip text, the stripped body (and the IR decoded again from it) is what gets redacted, recorded and sent on every hop. A refused request still leaves a failed row in traffic. Findings are reported on the request id. - Compaction requests (`/v1/responses/compact`, Codex `/backend-api/codex/responses/compact`) are screened like generation requests: they carry the whole conversation and run a model. Token counts are not screened: no model runs, and screening them would record the same finding twice and could refuse the client's count. - WebSocket frames: decodable `response.create` frames are screened and stripped like HTTP requests; other frames go through code point rules. - The output limit is removed from every path. - Fix: the excerpt of a flagged tool call is masked with the same redaction as stored bodies before it goes into `tool_call_flagged`. A secret restored from a placeholder used to reach the event, the security log and the system notification in the clear. Control API (protocol 33, request store SCHEMA 24, history is reset): - `Guard` is `redact | inspect_tools | content`; `SecurityDetail` and `SecurityView` have those three; `SetSecurityLimit` is removed. - Rule views, the test endpoint and their types come from tw-guard (`tw_guard::view`, `tw_guard::trial`); tests return `output` and `refused`. `SecurityRuleView` has `label`; `RuleAction` has `strip`; `ContentMatch` has `codepoints`; `CustomRuleSave` has `label`. - `content_matched` carries `match`, `action`, `outcome` (`recorded | stripped | blocked`) and `revealed`; `hidden_text_found` and `output_limited` are gone. `SecurityOutcome` gains `stripped`; the log entry gains `match` and `revealed`; summary counts are `secrets, secrets_replaced, tool_calls, tool_calls_cut, content, content_blocked, content_stripped`. Message codes added: config.rule_codepoints_bad, config.rule_label_bad, gw.content.refused_invisible_message, gw.content.refused_invisible_tool_result, security.bad_codepoints, security.bad_label, security.content_action_unknown, security.pattern_empty, security.unknown_guard. Removed: config.output_limit_range, gw.hidden_text.refused_message, gw.hidden_text.refused_tool_result, gw.output_limit.cut, gw.output_limit.withheld, security.guard_unknown, security.limit_range, security.no_custom_rules, security.no_limit, security.nothing_to_test, security.unknown_content_action. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say why legacy completions are not content-screened /v1/completions and /v1/complete carry one prompt string, with no way to tell the caller's text from a tool result, and today they mostly serve editor code completion, where there are no tool results to carry an injection. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say which rule a cut tool call matched, not who produced it (#269) The tool-call guard judges the answer as the client would receive it, and a script plugin with the tool-call permission can produce or modify tool calls. "The call returned by upstream X" is no longer always true, so the sentences now say the answer contained the call. New message codes (the old ones are removed): - gw.toolcall.cut -> gw.toolcall.response_cut - gw.toolcall.blocked -> gw.toolcall.response_withheld - gw.ws.toolcall_cut -> gw.toolcall.connection_cut All three take the same arguments: upstream, tool, rule, name, why (the WebSocket code used to call the reason `detail`). The upstream stays an argument, as it stays in tool_call_flagged and in the log fields. A WebSocket connection cut for a tool call now tells the client the same coded sentence its ending records, as a frame the content filter refuses already does, instead of a separate hard-coded line. The log lines no longer name the upstream as the source either. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * README: three protections, and the content filter removes hidden instructions The output limit is gone and hidden characters are content filter rules, so the highlights and the crate table describe three protections: outbound redaction, tool-call inspection and the content filter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Builds on guard-unify stage 1. Targets
feat/guard-unify; please do not merge from my side.Why
Outbound redaction replaces secrets with placeholders and restores them in the answer; that is unchanged. Tool calls are judged on the final, restored call the client will execute. The gap: a malicious upstream can write a tool call that, after restoration, sends a credential to an unknown host (e.g. a curl with the key in the query). Two built-in rules in the "dangerous" group close it.
Rules
secret-to-unknown-host— Send a credential to an unknown host. High; cut on enforce. Fires when a tool call makes an http(s) request and its arguments carry a value the outbound-redaction detector (tw_guard::redact) recognizes as a credential (Kind::ApiKeysorKind::PrivateKeys), while the destination is neither local (loopback,localhost,*.localhost) nor the credential's own provider. JWTs and connection strings are intentionally not treated as credentials here (bearer tokens travel to many hosts; too noisy to cut).upload-file-to-host— Upload a local file to an external host. Medium; record only. Fires when a tool call uploads a local file's contents to a non-local host (curl -T/--data @file/-F field=@file/--upload-file/--post-file). Common in development, so it only records by default.Representation
RuleSpecgains an optionalcheck: Option<String>naming a code-backed check. Code-backed rules carrypattern: ''indata/rules.yaml.tools::rules::Rulegainscheck: Option<Check>. Code-backed rules get a never-matchingre([^\s\S]) and are left out ofscan_rules(), so the client-config scanner and anything still readingrule.redirectly (Litetw-scan, the Enterprise test endpoint) simply never match them.Rule::find(args) -> Option<Found>handles both kinds. The wall, the shared trial and the desktop (interim) test endpoint call it instead ofrule.re.find.tw_guard::tools::net(Check, the provider map, URL/host parsing). The match range is thescheme://hostonly, so the excerpt never carries a credential from the query string.Matcher::Builtin { check }so the security page can list the rules.Scope / provider map
Provider hosts (suffix match, dotted boundary), keyed by redaction rule id: Anthropic→anthropic.com; OpenAI→openai.com; GitHub→github.com, githubusercontent.com; Slack→slack.com; AWS→amazonaws.com; Google→googleapis.com, google.com; GitLab→gitlab.com; Stripe→stripe.com; npm→npmjs.org, npmjs.com; DigitalOcean→digitalocean.com; SendGrid→sendgrid.com. Private keys have no provider (any non-local host fires).
Not in scope
No shell/listener rules, per the brief.
relay.rs,ws.rs, the cut message and excerpt masking are untouched (stage 2 owns those).Checks
cargo fmt --all --check,cargo clippy --workspace --all-targets -- -D warnings,cargo test --workspace, and thetssteps (cargo test -p tw-api --features ts,export_ts) all pass locally.docs/config.md+.zh-CN.mdregenerated.🤖 Generated with Claude Code