Unify the request guards with thinkwatch-core's rule model - #72
Merged
Merged
Conversation
Outbound redaction, tool-call inspection and the content filter now run
on the shared model in tw-guard: the policy shape, built-in catalogs,
validation, rule view and trial, the content `screen` and the redaction
`flow`. The enterprise copies of each are gone.
- Settings: `security.redact`, `security.inspect_tools` and
`security.content` hold one policy each, validated by core and seeded
as the factory policy (observe). The old keys and
`models.output_guardrails` are converted once at boot, in one
transaction, keeping what they did, then removed.
- Content filter: rules refuse (403), strip the matched text (the
request continues as the stripped body, decoded again) or record;
code-point rules; hidden characters are built-in rules now.
- Redaction: the whole request body is searched, every hop goes out
through the same ledger, placeholders are `<<TW_LABEL_n>>`.
- Tool calls are judged as the client receives them, restored.
- One audit event per hit; excerpts and refusal quotes are masked with
the redaction rules before they are written.
- Models: `max_output_tokens` caps the output a request may ask for,
forwarded or converted, replacing the output length guardrail.
- Console API: `GET /api/admin/security`,
`POST /api/admin/security/{guard}/test`; the five old content-filter,
PII and tool-inspection endpoints are removed. Writing a guard policy
takes `pii_redactor:write` or `content_filter:write`.
Core is pinned to `feat/guard-unify` by rev until it is released.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… model
The security page shows the three request guards - outbound redaction,
tool-call inspection and the content filter - one tab each, in the shape
the shared guard model gives both products:
- A mode card per guard: Off / Observe / a third mode named for what it
does there (Replace, Cut off, Enforce), what the current mode does,
and the cost of the third mode, said before anyone switches to it.
- A rule table grouped as the server lists it (content: hidden
characters, instruction override, identity and prompts, Chinese
phrasing, custom) with what each rule matches, what it does
(redaction: the placeholder a hit becomes) and a switch.
- Dialogs to create and edit custom rules (content rules match by text,
regex or code points, and refuse, delete or record; redaction rules
name their placeholder, SECRET unless changed), to view a built-in
rule and change its action, and to test a sample against a guard:
the hits, the text as the third mode would send it, and whether the
content filter would refuse the request.
Every change writes the guard's whole policy object to its
security.redact / security.inspect_tools / security.content settings
key, built from the rule view GET /api/admin/security returns; samples
are tried through POST /api/admin/security/{guard}/test. Mode and
switch changes show at once and offer an undo. Writes stay behind the
existing permissions: pii_redactor:* for redaction, content_filter:*
for the other two.
The hidden-character card, the content-filter presets and the old test
sandbox are gone: hidden characters are content rules now.
The model editor's output guardrails give way to "Max output tokens"
(blank for no limit), sent as max_output_tokens; the model drawer
shows it.
The types for the new endpoints are provisional (src/lib/security-types.ts)
until the backend's OpenAPI schema carries them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checked against tw-guard's view.rs and trial.rs at the pinned core
commit and the handlers in this branch:
- Tool-call rules implemented in code carry `{ kind: "builtin", check }`
instead of a regex. The two shipped ones - secret-to-unknown-host
(cuts) and upload-file-to-host (records) - get names and reasons in
both languages, and the match column and the rule dialog say what the
check looks for. They have no pattern, so they cannot be copied as a
custom rule. The column header is "Match" for every guard now that
not every tool rule is a regex.
- A single rule is tried with the action chosen in its dialog (the test
request's new `action`, never sent for redaction), so the marks, the
text the third mode would send and the refusal all follow the choice;
the dialogs show what the server answers instead of guessing.
- The page needs `settings:read`, as GET /api/admin/security does;
writes and tests keep the guards' own permissions.
- The types say where they come from and how the server fills them
(`why` absent rather than null, `trial` as the rule of a tried
pattern, `output` null when the request would be refused).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README (both languages) still described PII patterns, hidden character detection set to warn and a content filter that blocks, warns or logs. It now describes the three guards as they ship: three modes each with the third named for what it does, every rule visible and switchable on the security page, redaction over the whole request with `<<TW_LABEL_n>>` placeholders, content rules that match phrases, regexes or code points and refuse, delete or record, hidden characters as content rules, the two built-in exfiltration rules for tool calls, and a model's maximum output tokens in place of the output length guardrail. The web README lists the security page and corrects what the Settings security tab holds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ThinkWatch-Core#268 is merged (72954bf) and its branch is gone, so the layer-one crates follow main by rev until the release tag exists. The layer-one code is the same as at d86320b. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The guard-unify work is released in ThinkWatch-Core v0.58.0, so the layer-one crates move from the main commit (72954bf) to the tag. Their code is unchanged between the two. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first v0.58.0 tag (d691c20) was deleted unreleased; the tag now points at c2a7bc6. The layer-one crates are the same in both. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The enterprise backend half of the guard unification: outbound redaction, tool-call inspection and the content filter now run on thinkwatch-core's shared rule model (
tw_guard::policy,view,trial,content::screen,redact::flow), the same one the desktop gateway uses. The console side is onguard-unify-weband is merged separately.Pinned to ThinkWatch-Core
v0.58.0(c2a7bc6), the release that carries ThinkWatchProject/ThinkWatch-Core#268; the layer-one crates move from v0.55.0.Settings
tw_guard::policydefines:security.redact,security.inspect_tools,security.content. Saved throughPATCH /api/admin/settings, checked by core's own validation (400 with its message). Seeds ship{}(the factory policy: observe).security.content_filter_patterns,security.pii_redactor_patterns,security.hidden_text,security.tool_inspection) andmodels.output_guardrailsare converted at boot, in one transaction that also deletes them, under an advisory lock (common::guard_policy::legacy). The conversion keeps behaviour: content rules identical to a built-in rule become that rule; a list with rules runs in enforce with unlisted built-ins switched off;hidden_textmaps ontounicode-tags/bidi-controls(blockwith an empty list still enforces); the four seeded PII patterns become built-in rules, the rest custom rules labelled by their old prefix; a byte cap of N becomesceil(N / 4)output tokens. Unit tests cover every mapping; an integration test runs it against a real database twice.Gateway
content::screenon the caller's own bytes; refuse (403, masked quote), strip (the request continues as the stripped body, decoded again — redaction, every hop and the audit row see it), or record. The Responses WebSocket goes through the same pipeline.flow::lookon the whole request body,flow::replaceon every hop with the same ledger, restoration by the shared engine;<<TW_LABEL_n>>placeholders.pii_redactor.rsand its own traversal are gone.ToolPolicy; calls are judged as the client receives them — converted and with placeholders restored — which is what core's newsecret-to-unknown-hostrule needs.gateway.content_{flagged,stripped,blocked},gateway.redaction_{flagged,replaced},gateway.tool_call_{flagged,blocked}. Excerpts (and the content refusal quote) are masked with the redaction rules first.output_guardrails.rsis gone. A model'smax_output_tokenslowers a larger ask and fills a missing one withtw_dialect::params::cap_max_output_tokens, forwarded or converted.Console API
GET /api/admin/security→tw_guard::view::detail(settings:read).POST /api/admin/security/{guard}/test→tw_guard::trial::run(pii_redactor:read/content_filter:read).security.redacttakespii_redactor:write;security.contentandsecurity.inspect_toolstakecontent_filter:write./api/admin/settings/content-filter/test,/content-filter/presets,/pii-redactor/test,/tool-inspection/rules,/tool-inspection/test.max_output_tokens(1–2147483647,nullclears) replacesoutput_guardrails.CHANGELOG has the upgrade notes under Unreleased.
Tests
cargo fmt --check,cargo clippy --workspace --all-targets -D warningsand the unit tests pass locally; the integration suite passed in full locally (379 tests) and runs in CI on every push. The web side came in fromguard-unify-web, and dev was merged in for the clippy fix (#73).🤖 Generated with Claude Code