Linear-time redaction for requests full of secrets, a per-request report cap, and output-cap edge cases - #279
Merged
Merged
Conversation
A request holding tens of thousands of distinct secrets (an exported credentials list, a log) made every redaction step quadratic: - findings merged each hit by searching the ones already merged; - hits on our own placeholders and on base64 payloads were dropped by comparing every hit with every span; - apply replaced in place from the back, moving the whole tail once per hit; - restoring, one-shot and streamed, searched the whole text once per placeholder in the ledger, and the stream's hold-back check compared the tail with every placeholder on every chunk. Now findings merge through a hash map, the overlap filters sort the spans and binary-search a running maximum of their ends, apply copies the text once from front to back, and restoring scans the text once, looking up each candidate placeholder by the few lengths the ledger issued. The stream restorer keeps the ledger sorted, shared by every lane of a stream, so a prefix check is a binary search. Results are byte for byte what they were: each change has a test that runs the old algorithm next to the new one. The one difference is a pathological restore where an original value itself contains placeholder text: it is no longer substituted again, which the old code did or not depending on hash map order. Measured in release mode, 40,000 distinct AWS access key IDs in an 840 KB request: look 2108 ms -> 21 ms, restoring the echoed answer 6804 ms -> 5 ms, streamed restore 6768 ms -> 30 ms. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The security log has one row per value a request carried, so a request with tens of thousands of distinct secrets filled it with that many rows and made one event several megabytes. A request now reports the first 100 distinct values, in the order they first appear. Every value is still replaced and restored: replacement follows the ledger, not the report. The cap is per request: when a plugin rewrites the request, the values it added are reported only up to what the first report left. Each frame a WebSocket client sends is a request of its own. tw-guard still returns every distinct value (findings), so a caller that aggregates instead of listing them, like the enterprise gateway's audit, has the full count. more_found, which picks out the values a plugin added, compares through a hash set instead of every pair. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The plugin bridge searched every string handed to a plugin once per value in the request's ledger, and on every streamed piece compared every prefix of every value with the buffer's tail. With tens of thousands of values both grew with values x text. It now finds every occurrence of every known value in one rolling-hash pass per distinct value length (each hash hit checked byte for byte), then keeps the occurrences the old value-by-value replace kept: longer values first, and of one value the leftmost non-overlapping ones. The hold-back check binary-searches the values sorted by bytes. Tests run the old loops next to the new code on random text, with values that are prefixes, suffixes and overlaps of each other. The only difference is a value that happens to be part of a placeholder just put in (a password "1"): the old loop corrupted that placeholder, this one does not, since it searches the original text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…et, keep thinking valid
Three edge cases of the enterprise gateway's per-model output cap. The
desktop's routing-rule setter (set_max_output_tokens) is unchanged.
- A Chat request naming both max_completion_tokens and max_tokens had
only the first checked, so {"max_completion_tokens": 100,
"max_tokens": 9999} went out unchanged and an upstream reading only
max_tokens was not capped. Each field written is now lowered to the
cap on its own; one already under it is left alone, and null counts
as not written.
- A Chat request naming neither got max_tokens, which OpenAI's
reasoning models refuse with 400. The function takes `official`, the
same flag as Target::official (tw_dialect::official::is_official_host
of the upstream): an official endpoint gets max_completion_tokens,
any other max_tokens, as the encoder does.
- An Anthropic request with thinking enabled must keep budget_tokens
below max_tokens, since thinking tokens count toward max_tokens.
When the cap lowers max_tokens to or below the budget, the budget
goes to cap - 1 if the cap is above 1024 (the smallest budget), and
thinking is removed otherwise. The same applies to Claude behind
Converse (additionalModelRequestFields.thinking).
Signature: cap_max_output_tokens(dialect, body, cap, official).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Found in the regression review before enterprise 3.0.0. Layer one (tw-guard, tw-dialect) is shared, so the desktop gateway had the same problems.
1. Outbound redaction is linear in the number of distinct values (tw-guard)
A request with tens of thousands of distinct secrets (an exported credentials list, a log) made every step quadratic. Each spot found, and the fix:
rules::findingsflow::hitsoutside)replace::applyreplace_rangefrom the back: the whole tail moved once per hitreplace::restore/restore_json<<, look up the few placeholder lengths the ledger issued (Ledgernow tracks them)stream::RestorerRestorer::fresh); the prefix check is a binary searchResults are byte for byte what they were. Every change has a test that runs the old algorithm next to the new one on generated input: findings order, numbering, dedup; overlap filtering with unsorted, nested, empty spans; apply with repeated values, labels and a seeded ledger; restore with adjacent placeholders,
<<<, half-written and unknown placeholders; hold-back points and chunked streaming. One case differs, and it was undefined before. If a restored value itself contains placeholder text, the old code might replace it again, depending on hash map order. The new code never does.tests/scale.rsruns the whole round trip with 10,000 and 40,000 distinct values: look, replace, whole restore, chunked restore and an Anthropic SSE stream. It checks every result and asserts that the 40k run takes less than 10× the 10k run, using the fastest of three runs each. Linear is about 4×, quadratic 16×. Locally the ratio is 4.3 in debug. Against the old sources, the test did not finish in 10 minutes.2. At most 100 reported values per request (tw-gateway)
The security log has one row per value (
SecretsFound.items).guard::itemsnow reports the first 100 distinct values in order of first appearance (REPORTED_MAX). Every value is still replaced and restored: replacement follows the ledger, not the report. The cap is per request. When a plugin rewrites the request, the second event only reports the plugin-added values that still fit under 100. Each WebSocketresponse.createframe is a request of its own. tw-guard still returns every distinct value fromlook/findings, so the enterprise audit can aggregate with the full count (findings.len()).more_foundnow compares through a hash set instead of every pair.3. Plugin bridge (tw-gateway)
Bridge::hidesearched each string given to a plugin once per ledger value.hold_fromcompared every prefix of every value on every streamed piece. Both grew with values × text.hidenow does one rolling-hash pass per distinct value length, checking each hash hit byte for byte. It keeps the occurrences the old longest-first, value-by-valuereplacekept.hold_frombinary-searches the values sorted by bytes. There are comparison tests against the old loops and a scale test (5k vs 20k values: 4.1× locally). One case differs: when a value is part of a placeholder just inserted (a password1), the old loop corrupted that placeholder.4.
cap_max_output_tokensedge cases (tw-dialect, enterprise only)New signature:
cap_max_output_tokens(dialect, body, cap, official: bool) -> bool.officialmeans the same asTarget::official(official::is_official_hostof the upstream).max_completion_tokensandmax_tokens: each is lowered to the cap on its own. A field already under the cap is left alone, andnullcounts as not written.max_completion_tokens, any othermax_tokens. OpenAI reasoning models refusemax_tokenswith a 400.type: enabledwithbudget_tokens): thinking tokens count towardmax_tokens, so the budget must stay below it. When the cap lowersmax_tokensto or below the budget, the budget becomescap - 1if the cap is above 1024. Otherwisethinkingis removed. The same applies to Claude behind Converse (additionalModelRequestFields.thinking). A request whose ownmax_tokenswas already at or below the cap is not touched.set_max_output_tokens, used by the desktop's routing rules, is unchanged. A test pins it.Each case is tested across Chat, Anthropic, Responses and Gemini, plus Converse.
Measurements (release, Apple Silicon, one 840 KB user message of distinct
AKIA…keys)flow::lookflow::replacerestore(answer echoes all)At 500,000 values the old code was quadratic: by extrapolation, minutes for
lookand over 15 minutes for restore.Compatibility
UPDATE_CONFIG_DOCS/UPDATE_MSG_CODES.flow::off_placeholders,stream::Restorer::fresh).replace::applynow states its existing contract (hits sorted and disjoint, asscanreturns) and checks it with adebug_assert!.cap_max_output_tokenssignature change breaks enterprisedev. ThinkWatch branchfix/redact-scale(same name, fromdev) passesfalseat the current pre-routing call, which keeps today'smax_tokens, so the enterprise job builds against this PR. The per-hop cap infix/guard-review(cap_output(.., _official)) can pass the real value once it pins this core.release-notes/0.59.0.md(redaction speed and the 100-value cap). Thecap_max_output_tokenschanges are enterprise-only and not in the desktop notes.Checked locally:
cargo fmt --all --check,cargo clippy --workspace --all-targets -- -D warnings,cargo test --workspace(proxy unset),python3 scripts/release_notes_test.py.🤖 Generated with Claude Code