Skip to content

Linear-time redaction for requests full of secrets, a per-request report cap, and output-cap edge cases - #279

Merged
fylorn merged 5 commits into
mainfrom
fix/redact-scale
Oct 3, 2026
Merged

fylorn merged 5 commits into
mainfrom
fix/redact-scale

Conversation

@fylorn

@fylorn fylorn commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Found in the regression review before enterprise 3.0.0. Layer one (tw-guard, tw-dialect) is shared, so the desktop gateway had the same problems.

1. Outbound redaction is linear in the number of distinct values (tw-guard)

A request with tens of thousands of distinct secrets (an exported credentials list, a log) made every step quadratic. Each spot found, and the fix:

Where Before Now
rules::findings merged each hit by searching the findings already merged hash map keyed on (rule, value)
flow::hits dropped hits on our placeholders / base64 payloads by comparing every hit with every span spans sorted once, running max of their ends, one binary search per hit (outside)
replace::apply replace_range from the back: the whole tail moved once per hit one front-to-back copy
replace::restore / restore_json searched the whole text once per placeholder in the ledger one scan; at each <<, look up the few placeholder lengths the ledger issued (Ledger now tracks them)
stream::Restorer same restore per chunk, plus a hold-back check comparing the tail with every placeholder per chunk the ledger sorted once and shared by every lane of a stream (Restorer::fresh); the prefix check is a binary search

Results are byte for byte what they were. Every change has a test that runs the old algorithm next to the new one on generated input: findings order, numbering, dedup; overlap filtering with unsorted, nested, empty spans; apply with repeated values, labels and a seeded ledger; restore with adjacent placeholders, <<<, half-written and unknown placeholders; hold-back points and chunked streaming. One case differs, and it was undefined before. If a restored value itself contains placeholder text, the old code might replace it again, depending on hash map order. The new code never does.

tests/scale.rs runs the whole round trip with 10,000 and 40,000 distinct values: look, replace, whole restore, chunked restore and an Anthropic SSE stream. It checks every result and asserts that the 40k run takes less than 10× the 10k run, using the fastest of three runs each. Linear is about 4×, quadratic 16×. Locally the ratio is 4.3 in debug. Against the old sources, the test did not finish in 10 minutes.

2. At most 100 reported values per request (tw-gateway)

The security log has one row per value (SecretsFound.items). guard::items now reports the first 100 distinct values in order of first appearance (REPORTED_MAX). Every value is still replaced and restored: replacement follows the ledger, not the report. The cap is per request. When a plugin rewrites the request, the second event only reports the plugin-added values that still fit under 100. Each WebSocket response.create frame is a request of its own. tw-guard still returns every distinct value from look/findings, so the enterprise audit can aggregate with the full count (findings.len()). more_found now compares through a hash set instead of every pair.

3. Plugin bridge (tw-gateway)

Bridge::hide searched each string given to a plugin once per ledger value. hold_from compared every prefix of every value on every streamed piece. Both grew with values × text.

hide now does one rolling-hash pass per distinct value length, checking each hash hit byte for byte. It keeps the occurrences the old longest-first, value-by-value replace kept. hold_from binary-searches the values sorted by bytes. There are comparison tests against the old loops and a scale test (5k vs 20k values: 4.1× locally). One case differs: when a value is part of a placeholder just inserted (a password 1), the old loop corrupted that placeholder.

4. cap_max_output_tokens edge cases (tw-dialect, enterprise only)

New signature: cap_max_output_tokens(dialect, body, cap, official: bool) -> bool. official means the same as Target::official (official::is_official_host of the upstream).

  • Chat with both max_completion_tokens and max_tokens: each is lowered to the cap on its own. A field already under the cap is left alone, and null counts as not written.
  • Chat with neither: an official endpoint gets max_completion_tokens, any other max_tokens. OpenAI reasoning models refuse max_tokens with a 400.
  • Anthropic thinking (type: enabled with budget_tokens): thinking tokens count toward max_tokens, so the budget must stay below it. When the cap lowers max_tokens to or below the budget, the budget becomes cap - 1 if the cap is above 1024. Otherwise thinking is removed. The same applies to Claude behind Converse (additionalModelRequestFields.thinking). A request whose own max_tokens was already at or below the cap is not touched.
  • set_max_output_tokens, used by the desktop's routing rules, is unchanged. A test pins it.

Each case is tested across Chat, Anthropic, Responses and Gemini, plus Converse.

Measurements (release, Apple Silicon, one 840 KB user message of distinct AKIA… keys)

distinct values flow::look flow::replace restore (answer echoes all) streamed restore (64 B chunks)
10,000, before 149 ms 14 ms 392 ms 353 ms
10,000, after 8 ms 4 ms 1 ms 8 ms
40,000, before 2108 ms 250 ms 6804 ms 6768 ms
40,000, after 21 ms 12 ms 6 ms 33 ms
500,000 (10.5 MB), after 431 ms 292 ms 102 ms 666 ms

At 500,000 values the old code was quadratic: by extrapolation, minutes for look and over 15 minutes for restore.

Compatibility

  • No protocol change (tw-api untouched), no message-code change, no config change. Tests pass without UPDATE_CONFIG_DOCS / UPDATE_MSG_CODES.
  • tw-guard: additions only (flow::off_placeholders, stream::Restorer::fresh). replace::apply now states its existing contract (hits sorted and disjoint, as scan returns) and checks it with a debug_assert!.
  • tw-dialect: the cap_max_output_tokens signature change breaks enterprise dev. ThinkWatch branch fix/redact-scale (same name, from dev) passes false at the current pre-routing call, which keeps today's max_tokens, so the enterprise job builds against this PR. The per-hop cap in fix/guard-review (cap_output(.., _official)) can pass the real value once it pins this core.
  • Release notes: one section in release-notes/0.59.0.md (redaction speed and the 100-value cap). The cap_max_output_tokens changes are enterprise-only and not in the desktop notes.

Checked locally: cargo fmt --all --check, cargo clippy --workspace --all-targets -- -D warnings, cargo test --workspace (proxy unset), python3 scripts/release_notes_test.py.

🤖 Generated with Claude Code

fylorn and others added 5 commits October 3, 2026 17:50
A request holding tens of thousands of distinct secrets (an exported
credentials list, a log) made every redaction step quadratic:

- findings merged each hit by searching the ones already merged;
- hits on our own placeholders and on base64 payloads were dropped by
  comparing every hit with every span;
- apply replaced in place from the back, moving the whole tail once
  per hit;
- restoring, one-shot and streamed, searched the whole text once per
  placeholder in the ledger, and the stream's hold-back check compared
  the tail with every placeholder on every chunk.

Now findings merge through a hash map, the overlap filters sort the
spans and binary-search a running maximum of their ends, apply copies
the text once from front to back, and restoring scans the text once,
looking up each candidate placeholder by the few lengths the ledger
issued. The stream restorer keeps the ledger sorted, shared by every
lane of a stream, so a prefix check is a binary search.

Results are byte for byte what they were: each change has a test that
runs the old algorithm next to the new one. The one difference is a
pathological restore where an original value itself contains placeholder
text: it is no longer substituted again, which the old code did or not
depending on hash map order.

Measured in release mode, 40,000 distinct AWS access key IDs in an
840 KB request: look 2108 ms -> 21 ms, restoring the echoed answer
6804 ms -> 5 ms, streamed restore 6768 ms -> 30 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The security log has one row per value a request carried, so a request
with tens of thousands of distinct secrets filled it with that many rows
and made one event several megabytes. A request now reports the first
100 distinct values, in the order they first appear. Every value is
still replaced and restored: replacement follows the ledger, not the
report.

The cap is per request: when a plugin rewrites the request, the values
it added are reported only up to what the first report left. Each frame
a WebSocket client sends is a request of its own.

tw-guard still returns every distinct value (findings), so a caller that
aggregates instead of listing them, like the enterprise gateway's audit,
has the full count.

more_found, which picks out the values a plugin added, compares through
a hash set instead of every pair.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The plugin bridge searched every string handed to a plugin once per
value in the request's ledger, and on every streamed piece compared
every prefix of every value with the buffer's tail. With tens of
thousands of values both grew with values x text.

It now finds every occurrence of every known value in one rolling-hash
pass per distinct value length (each hash hit checked byte for byte),
then keeps the occurrences the old value-by-value replace kept: longer
values first, and of one value the leftmost non-overlapping ones. The
hold-back check binary-searches the values sorted by bytes.

Tests run the old loops next to the new code on random text, with
values that are prefixes, suffixes and overlaps of each other. The only
difference is a value that happens to be part of a placeholder just put
in (a password "1"): the old loop corrupted that placeholder, this one
does not, since it searches the original text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…et, keep thinking valid

Three edge cases of the enterprise gateway's per-model output cap. The
desktop's routing-rule setter (set_max_output_tokens) is unchanged.

- A Chat request naming both max_completion_tokens and max_tokens had
  only the first checked, so {"max_completion_tokens": 100,
  "max_tokens": 9999} went out unchanged and an upstream reading only
  max_tokens was not capped. Each field written is now lowered to the
  cap on its own; one already under it is left alone, and null counts
  as not written.
- A Chat request naming neither got max_tokens, which OpenAI's
  reasoning models refuse with 400. The function takes `official`, the
  same flag as Target::official (tw_dialect::official::is_official_host
  of the upstream): an official endpoint gets max_completion_tokens,
  any other max_tokens, as the encoder does.
- An Anthropic request with thinking enabled must keep budget_tokens
  below max_tokens, since thinking tokens count toward max_tokens.
  When the cap lowers max_tokens to or below the budget, the budget
  goes to cap - 1 if the cap is above 1024 (the smallest budget), and
  thinking is removed otherwise. The same applies to Claude behind
  Converse (additionalModelRequestFields.thinking).

Signature: cap_max_output_tokens(dialect, body, cap, official).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit ef10a8b into main Oct 3, 2026
4 checks passed
@fylorn
fylorn deleted the fix/redact-scale branch October 3, 2026 10:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant