Repository navigation
feat(opf): scan in the background so git push never waits on the privacy filter - #2631
peyton-alt wants to merge 23 commits into
Conversation
BatchBytesWithPrivacyFilter ran the model and applied its spans in one call, so whoever needed redacted output had to wait for inference. Split it: - ScanBlobsWithPrivacyFilter runs OPF over blobs with no cache entry and stores, per blob, the spans for each prose leaf, keyed by the blob's object hash and the category set. Entries hold leaf hashes and offsets, no text. - ApplyCachedPrivacyFilter redacts from the cache alone and returns ErrOPFScanPending unless every blob and every leaf is covered. - BatchBytesWithPrivacyFilter keeps its behavior and shares the same leaf collection and scan code. Model calls are split into chunks of at most 1 MiB of leaf text, each with a deadline scaled to its size (floor 30s, ceiling 3h) instead of a fixed 30s, and the batch input ceiling rises to 256 MiB. The throughput and bug investigation docs these numbers come from are included. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SPBMJ9H9KF399NA898R42W
Implements redact.OPFSpanCache as one small JSON file per blob under entire-opf-cache/, written atomically through the shared git-common-dir root. Keys must be SHA-256 digests. Entries older than 30 days can be pruned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SPEEGMR0DFFFXQAJJ85VWW
Both rewrites (git-refs per queued ref, git-branch over the unpushed v1 chain) take a mode. opfScanThenApply scans anything the span cache lacks and then applies it; the exported RewriteQueuedCheckpointRefsWithOPF and RewriteUnpushedV1WithOPF keep that synchronous behavior. opfApplyCachedOnly never runs the model and reports unscanned content as OPFScanPendingError, leaving those refs untouched while covered siblings are rewritten. The cache is resolved from the repository being rewritten, not the process's working directory. Collected blobs now carry their object IDs. The OPF batch cap is a pathological-input backstop again (128 MiB prose, 256 MiB raw), since the model no longer runs inside git push. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SPSCTZEG1M9KD6357HWC8D
The pre-push hook no longer runs the OPF model. It rewrites checkpoints from the span cache and holds back anything not scanned yet: on git-refs those refs stay queued while scanned siblings ship, on git-branch v1 is not pushed this time. Either way the user's push completes and a detached `entire __opf_scan <remote>` worker is started. The worker takes a per-repo lock for its whole run, collects the checkpoint commits still lacking the OPF trailer on the primary backend (queued refs on git-refs, the unpushed v1 chain on git-branch), scans what the cache lacks, and then runs the same pre-push delivery for the remote the spawning push named, with that push's OPF decision and non-interactive git. Units over a cap are skipped, the loop repeats only while new work appears, and a delivery that fails leaves everything queued for the next push. The explicit migration push still scans synchronously, since it promises an immediate push. Also carries two fixes the worker depends on: SpawnDetached gives children the null device instead of a pipe (a pipe drained by the exiting parent killed the child on its first stderr write), and the spawn-throttle marker moves into a shared spawnmarker package. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SQFKMHPMBCMBFZZJFSD4G5
Pre-push printed that held refs "stay queued for the next push" while the worker was about to push them itself, and the non-interactive and prompt text still promised a ~30s scan during the push. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SQMZAXBRY6KDR8D79Z5GHJ
entire status (text and --json checkpoint_opf_pending) reports how many checkpoints are held until the background scan covers them, on both backends: queued refs without the trailer on git-refs, unpushed untrailered v1 commits on git-branch. The count reads local refs only and is omitted when OPF is off. When v1 is held for a scan, git-branch users see a tip, at most once every 30 days per repository, pointing at `entire doctor migrate-checkpoints`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SQRDXRVVA9WV1Y0VYAMXZ5
The security doc now says OPF never runs inside git push: pre-push holds unscanned checkpoints, the scan worker scans and delivers them, the span cache stores only hashes and offsets, and failures keep checkpoints held rather than aborting pushes. Cap values, timeout_seconds, the prompt text and the CI note match the new behavior. The OPF command-trust integration tests wait for the background worker to log that it finished before checking whether the payload ran, since that is now where the opf binary executes; otherwise the negative tests would pass vacuously. The worker logs that line on every exit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3SRQS5BCD884ER2E6ZX76KD
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 05e1e64. Configure here.
There was a problem hiding this comment.
Note
Copilot was unable to run its full agentic suite in this review.
Copilot review overview
Review effort: Lite
Findings: 1
Open (5)
On any lock acquisition error other thancontext.DeadlineExceeded, this returns `(noop,… · New The chunk sizing accounts forlen(opfBatchSeparator)per input, which only matches the real batch… · New The estimate uses integer truncation (batchedLen/1024), which rounds down and slightly… · New The interface doc says implementations only need to be safe 'from one process at a time', but this… · New The comment references a__opf_flushchild as the motivating example, but this PR introduces… · New
What changed in this PR
This PR refactors OpenAI Privacy Filter (OPF) handling to avoid running model inference on the git push hot path by introducing a persistent span cache plus a detached background scan worker, while also updating OPF batching/timeouts and exposing backlog visibility via entire status.
Changes:
- Add an OPF span cache (git-common-dir) with scan/apply split APIs and tests.
- Introduce a detached
__opf_scanworker that scans held checkpoints and then delivers them, while pre-push rewrites from cache only. - Make OPF shell-out deadlines size-adaptive and chunk large batches into bounded calls; update docs and status surfaces accordingly.
| File | Description |
|---|---|
| redact/opf_test.go | Updates timeout-default behavior tests; adds adaptive timeout + deadline wiring tests |
| redact/opf_cache_test.go | Adds tests for scan-then-apply caching behavior and safety properties |
| redact/opf_cache.go | Introduces OPFSpanCache interface + scan/apply cached flow |
| redact/opf.go | Adds size-adaptive timeouts; increases raw input cap; uses adaptive timeout in shell-out |
| redact/batch_test.go | Adds tests for chunking logic and bounded-call splitting behavior |
| redact/batch.go | Refactors leaf collection; adds chunked scanning; adds NamedBlob.ID and helpers |
| docs/security-and-privacy.md | Updates user-facing behavior/docs for background scanning, caps, and timeouts |
| docs/development/opf-throughput-findings.md | Adds benchmark findings motivating architectural changes |
| docs/development/opf-bug-investigation.md | Adds investigation write-up motivating new caps/flow |
| cmd/entire/cli/trail_context_cache.go | Switches to shared spawnmarker throttling helper |
| cmd/entire/cli/strategy/manual_commit_push_test.go | Updates per-ref delivery test to reflect cache-only pre-push and worker spawning |
| cmd/entire/cli/strategy/manual_commit_push.go | Holds unscanned checkpoints, spawns worker, rewrites from cache-only on pre-push |
| cmd/entire/cli/strategy/manual_commit_opf_scan_test.go | Adds tests for scan worker behavior, locking, and delivery |
| cmd/entire/cli/strategy/manual_commit_opf_scan.go | Adds detached worker implementation + backlog counting helper |
| cmd/entire/cli/strategy/manual_commit_opf_rewrite_test.go | Updates rewrite tests for cache-only mode and background-worker semantics |
| cmd/entire/cli/strategy/manual_commit_opf_rewrite.go | Adds rewrite modes, pending error, cache-backed redact path, and new caps |
| cmd/entire/cli/strategy/manual_commit_opf_refs.go | Adds mode-aware rewrite path for git-refs backend |
| cmd/entire/cli/strategy/manual_commit_opf_prompt_test.go | Updates prompt/progress messaging expectation |
| cmd/entire/cli/strategy/manual_commit_opf_prompt.go | Updates non-interactive progress line and prompt description |
| cmd/entire/cli/status_test.go | Adds OPF backlog visibility tests (text + JSON) |
| cmd/entire/cli/status.go | Adds OPF backlog computation and status rendering/JSON field |
| cmd/entire/cli/spawnmarker/spawnmarker.go | Introduces shared spawn marker throttling utility |
| cmd/entire/cli/settings/settings.go | Adds settings-level OPFEnabled helper (distinct from runtime-configured OPFEnabled) |
| cmd/entire/cli/session_sweep.go | Switches to spawnmarker throttling helper |
| cmd/entire/cli/root.go | Registers new hidden __opf_scan command |
| cmd/entire/cli/opf_scan_cmd_test.go | Tests __opf_scan wiring through root + file logging |
| cmd/entire/cli/opf_scan_cmd.go | Adds hidden __opf_scan command implementation |
| cmd/entire/cli/integration_test/opf_command_trust_test.go | Updates integration tests to wait for background worker completion |
| cmd/entire/cli/execx/spawn_detached_test.go | Adds test ensuring detached child stdio is null-device |
| cmd/entire/cli/execx/spawn_detached.go | Refactors spawn into testable detachedCommand; avoids pipes that can SIGPIPE child |
| cmd/entire/cli/checkpoint/opf_span_cache_test.go | Adds git-common-dir OPF span cache tests + pruning behavior |
| cmd/entire/cli/checkpoint/opf_span_cache.go | Adds on-disk OPF span cache implementation + pruning |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
When the OPF span cache could not be opened, pre-push reported the content as pending, so the user saw "being scanned in the background" on every push while the worker hit the same error and exited. It now withholds the checkpoints with OPFCacheUnavailableError, which names the cause and the remedy, and starts no worker. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3T3WEGGV1G4VW9FCFD8MGQ6
For arbitrary blob contents, applying cached results must equal the one-pass batch output, and an entry missing any single prose leaf must make the apply report pending rather than redact partially. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Entire-Checkpoint: 01M3T3WG1QND15NQ65TSXE0V20
…eyton/opf-scan-worker # Conflicts: # cmd/entire/cli/strategy/manual_commit_opf_refs.go # cmd/entire/cli/strategy/manual_commit_push.go
The scan worker collected every queued ref's blobs into one slice before scanning, so a long backlog of individually in-cap refs could grow its memory without bound. It now walks pending commits first and loads and scans their content in batches no larger than one ref's raw cap, releasing each batch before collecting the next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Spawn a worker whenever none holds the lock, instead of at most once a minute, and have an exiting worker re-check once after releasing the lock, so work held by a push in either window is not stranded. - Store scan results after each model call's worth of blobs, and deliver after each worker batch, so an interrupted scan keeps its progress and scanned checkpoints ship without waiting for the whole backlog. - Give the OPF rewrite's remote-tip temp ref a per-process name; a concurrent worker delivery and user push could delete each other's ref, which reads as "remote has no v1". - Skip units over the bootstrap limit in the worker, as the rewrite does, and clean up pushed shadow branches when v1 is held. - Drop "(runs in the background)" from the status line, which counts checkpoints no worker may be scanning, and document that the cache version must be bumped when the model changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4E83KBKGT7JNM0069G3KPD0
…tself On git-branch, holding back Entire's own v1 push does not stop the user's push from sending the branch (`git push --all`, `--mirror`, or an explicit refspec): git chooses what to send before the hook runs. The pre-push line now saves git's ref list, gives it to Entire, and replays it to whatever runs after (a chained hook or the next Husky line, e.g. git lfs pre-push). Under OPF, a push that carries v1 at anything other than the tip OPF just verified is refused. The same line ends the script when Entire fails. Previously a chained pre-push hook's exit status decided the push, so an OPF abort was lost. The signal is ENTIRE_PRE_PUSH_STDIN_REFS=1, not a flag, so an older binary run by the new script ignores it instead of failing every push. Installed scripts without it read as outdated and are reinstalled; an older hand-copied hook-manager line gets a warning when v1 is held. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4E9G4VHSPN4WY1JMRKQRSPN
The integration helper returned at the first "opf scan worker finished", so a worker spawned by a later push could still be writing to .git when the test's temp dir was removed. Log "finished" on every worker exit and wait until each spawn has a matching finish. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The outer-push guard only covered git-branch. On git-refs, `--mirror`, an explicit refspec, or `--all` with a v1 branch left from before a migration could still send checkpoint content OPF had not verified. The git-refs pre-push now refuses the user's push when it sends a checkpoint ref (v1 or refs/entire/checkpoints/*) whose commit lacks the OPF trailer, unless OPF is skipped for this push. On both backends a ref the remote already has at that commit is not counted as sent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ParsePrePushRefs capped stdin at 16 MiB, so a `--mirror` of a repository with many tags could truncate the list before the v1 line and silently skip the outer-push check. It now reads to the end and keeps only the lines that touch checkpoint refs, which bounds memory without a cap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4EBR89WEF9793069XD9F0QC
…eyton/opf-scan-worker # Conflicts: # cmd/entire/cli/status.go # cmd/entire/cli/strategy/manual_commit_push.go
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4EP93BENSX3HWM79AEBQ7TG
…eyton/opf-scan-worker
…nned ones The git-refs rewrite reported only its first error, so a ref over the bootstrap or size cap hid the pending signal from an unscanned sibling behind it. Pre-push then never started the scan worker, and the sibling stayed held on every push while the capped ref sat first in the queue. The rewrite now returns both, and the withheld warning prints the scan notice and the cap error. Also corrects two comments: the span cache is shared between the scan worker and pre-push, not used by one process at a time, and the detached child example names `__opf_scan` rather than the removed `__opf_flush`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4F6C6RP64X6BW8SF0QKVRSY
The previous fix joined the pending signal onto a cap error queued ahead of it, but a pending ref queued first still became the only error and the cap was dropped, so the user saw "scanning in the background" for a ref that will never ship. The rewrite now keeps pending apart from every other error and returns both, in either order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4F79SK3NGWYGRER6RXM7SVX
…rker - The pre-push line falls back to running Entire without the ref list when the temp file cannot be created or written, and an unreadable list is treated as unknown instead of failing the push. - A scan worker stops once the settings say OPF should never run. - `entire clean --all` removes the OPF span cache. - The span cache key includes the configured OPF command. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Entire-Checkpoint: 01M4GY0FJRX06GEEPFAA7YX2QN
…ng it Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>




https://entire.io/gh/entireio/cli/trails/1457
Stacked on #2533, which merges first; this PR depends on it.
Problem
OPF inference is slow, about 1.14 s per KB of prose on CPU, so one real session is minutes to hours of model time. The model ran inside the pre-push hook, which left two bad options. With the old 2 MiB cap, every real session was refused, so OPF could not ship real checkpoints at all. With the cap raised,
git pushwould wait for the whole scan, which ongit-branch(the default backend) could take hours.Result
git pushnever waits on the model, on either backend. Checkpoints still reach the remote only after OPF has scanned them.Scan and apply are split (
redact).ScanBlobsWithPrivacyFilterruns the model and stores results in a per-blob span cache.ApplyCachedPrivacyFilterredacts from the cache alone and reports anything not scanned yet. The cache lives inentire-opf-cache/in the git common directory. It stores only SHA-256 hashes of leaf text plus the offsets and labels OPF found, never text, keyed by blob hash and category set. Entries older than 30 days are pruned.Pre-push rewrites from the cache. Anything not scanned yet is held back: on git-refs that ref stays queued while scanned siblings ship, and on
git-branchv1 is not pushed this time. The user sees[entire] Checkpoints are being scanned by the OpenAI Privacy Filter in the background and will be pushed when it finishes.and their own push completes.Background worker (
entire __opf_scan <remote>, one per repo via a lock):If it cannot deliver, the checkpoints stay queued and the next push sends them from the cache without calling the model.
Capacity: model calls are split into chunks of at most 1 MiB, each with a deadline scaled to its size (30s floor, 3h ceiling). The caps are backstops for pathological input again: 128 MiB prose, 256 MiB raw.
Visibility:
entire status(text and--jsoncheckpoint_opf_pending) counts held checkpoints on both backends.git-branchusers with held checkpoints get a tip pointing atentire doctor migrate-checkpoints, at most once every 30 days.Unchanged: the explicit migration push still scans synchronously. No categories enabled, divergence and a Ctrl-C at the prompt behave as before.
Carried over from the earlier version of fix(strategy): dedup OPF byte cap, scope it per checkpoint ref, run it off git push's blocking path #2533:
SpawnDetachedgives children the null device instead of a pipe. With a pipe, the child was killed on its first stderr write oncegit pushreturned.Where to start reviewing
redact/opf_cache.goand theredact/batch.gorefactorstrategy/manual_commit_opf_scan.go: the worker, the spawn decision and the lockstrategy/manual_commit_push.go: pre-push gating on both backendsstrategy/manual_commit_opf_rewrite.go/_refs.go: the rewrite modesKnown gaps (not in this PR)
git push origin entire/checkpoints/v1is not blocked while v1 is unscanned. Pre-push does not read the refs git is pushing, and a rewrite inside the hook cannot change what git sends. This is the same on main.entire cleandoes not reclaimentire-opf-cache/yet. The worker's 30-day pruning bounds it.Validation
mise run check: lint, unit and integration tests, Vogon 56/56, Roger-Roger 4/4.FuzzApplyCachedPrivacyFilter: for arbitrary blob contents, cached apply equals the one-pass batch, and a missing leaf makes it report pending. Fuzzed for 90s (~610k inputs) with no failures.opfmodel and a local bare remote:git pushtook 0s with both held, and the worker delivered both about 20s later without another push. On the remote they are trailered, with 0 cleartext names and[REDACTED_PERSON]in their place.git-branch:git pushtook 1s with v1 held, and the worker pushed v1 about 15s later, redacted.🤖 Generated with Claude Code